# Synthetic fixture generators and decision evaluators

This is provenance for the Jev-Omni/Clef classification experiments. Generators create the test inputs; evaluators choose answers. No new media, ASR or model inference was run. The supplement covers all 35 accepted initial inputs plus five selected near-approach clips. Its JSON supplies each exact saved prompt/input text, requested seed, settings, selected attempt, source hashes and available observed media facts.

| Role | Exact model/checkpoint | Recorded revision and key settings |
| --- | --- | --- |
| Image generator | `ByteDance/SDXL-Lightning`, checkpoint `sdxl_lightning_4step_unet.safetensors`; base `stabilityai/stable-diffusion-xl-base-1.0` | Lightning `c9a24f48e1c025556787b0c58dd67a091ece2e44`; base `462165984030d82259a11f4367a4eed129e94a7b`. Four steps, guidance 0, EulerDiscreteScheduler with trailing timestep spacing; requested 1024×1024. Native PNG converted to exact evaluated JPEG at quality 90. |
| Video generator | `MiniMaxAI/MiniMax-H3` with Alibaba PDD8 acceleration | H3 `42ed227ee7df40d41602854ae760620d6eb651fe`; runtime source `cae69be584b8fc2daf0b03ec552710592f8738cb`. Recorded API setting: **9 steps**, flow shift 12 and audio flow shift 3, short edge 768, 16:9, four-second request, 24 fps. PDD8 is an acceleration label, not evidence of an eight-step API setting. Separate upstream PDD Hub revision is not recorded; available component hashes are in JSON. |
| Speech fixture generator | `ResembleAI/chatterbox-turbo` | `749d1c1a46eb10492095d68fbcf55691ccf137cd`; implementation source `5de7a54aa4e5e2baadb0182dde554908b48b85c2`. English, bundled default voice, no voice cloning. Original accepted files are full mono 24 kHz PCM16 WAVs. Seed is not exposed by the installed API; no deterministic regeneration claim. |
| Text fixture authoring | Original fictional Japanese emails | Authored before predictions; no generative model used. Exact email text appears in JSON and the existing inputs. |
| Decision evaluator | `akhilaaa3/Jev-Omni` | `c050d51354147985d13286cf4acf90f562f2c631`; one common loaded predictor across audio, images, text and video. Render technical model key `jev` as **Jev-Omni**. |
| Decision evaluator | `simonlehmann/clef-NVFP4` | `817ac58ad358f42489980ead62f8bf7cafc628c2`; community quantization of Cloudflare Clef, whose published backbone is Qwen/Qwen3.8-27B. No separately recorded upstream Qwen revision. Audio was not compared. |

Settings are separated into actual submitted request fields, recorded runtime/fixture configuration, and observed media. A saved prepared generation record alone is not an execution receipt. Accepted asset hashes and independent acceptance records bind the selected inputs. Seeds are the values in the submitted saved requests, including retries; they do not guarantee repeatability across hardware or library versions.

Initial native H3 clips have 107 frames and 4.458333-second recorded durations, while evaluated transports are the fixed first 96 frames at 24 fps, exactly four seconds. The selected near-approach inspection records the 107 native frames and exact four-second transport, but not native dimensions; the JSON preserves that unknown rather than presenting configured geometry as a fresh observation. Clef's eight and Jev-Omni's sixteen sampled frames are evaluator preprocessing, not generator steps.

The lineage is retained: five initial video slots remained missing after 20 attempts; three corkscrew generations were rejected before the desk-fan substitution; A01 selected attempt 01 rather than its earlier nonaccepted attempt. All ten strict-contact attempts failed contact acceptance. Their final attempt-1 clips were included only under the separate near-approach amendment, with ambiguous/not-needed references preserved. A prompt asking for contact is not evidence that the media contains contact.

For the Site worker, use this supplement with the public four-cohort export. Feature the four decision tasks in the main story and place detailed generation/selection lineage in a technical section. Chatterbox is fixture-generation provenance, not a separate speech-generation study. Original studies and response values remain unchanged; private runtime/auth/host details and execution logs are excluded.
