# Jev-Omni: one decision workflow across speech, images, email and video

Editorial synthesis for a Japanese YouTube video, based on completed research on 2026-10-02. Recommendations and possible application benefits below are editorial interpretation; numerical results are saved experiment observations. This document supersedes the earlier collision-centered upload narrative.

## Why recommend trying Jev-Omni?

Jev-Omni is an interesting model to try when an application must make decisions from several kinds of input. A spoken customer message, an object photo, an email and a short video can all become a task with a question and explicit choices. The appeal is a common decision workflow across video, images, text and audio.

That is the user's reason for recommending Jev-Omni here: useful coverage and a coherent way to prototype decisions. The completed study actually used one loaded Jev-Omni predictor for all four modalities. It did not measure production cost savings or prove that every input type shares an identical API: media used a Gateway route, while text called the common classifier through a direct route.

The practical demonstrations below show what this workflow can do on selected inputs. Possible uses such as customer-message triage, object quizzes, reply queues and clip review are application ideas, not production deployments evaluated in this research.

## Four practical demonstrations

### Speech: separate a complaint from an ordinary inquiry

An English customer says a parcel was marked delivered but never arrived, and requests delivery or a refund (A01). Another asks what time the shop opens on Saturdays (A06). These are two different decisions: a complaint asserting a service failure, and a prospective factual inquiry. Jev-Omni selected the frozen reference label for both examples in both choice orders.

The ten accepted Chatterbox Turbo speech clips contained five complaints and five ordinary inquiries. Jev-Omni matched **10/10 in each order**. The classifier received the full WAV, not a supplied ASR transcript. Independent saved ASR and semantic checks supported acceptance, with lexical uncertainty retained for A01. This demonstrates an audio-input decision task; it does not establish tone or emotion recognition. All ten original WAVs and the detailed mapping are supporting assets.

### Images: turn an object photograph into a quiz

A generated apple photograph was paired with Apple, Pear, Tomato and Peach (I01). Jev-Omni chose Apple in both orders. A perforated bowl with handles was paired with Mixing bowl, Colander, Fine-mesh sieve and Saucepan (I06). Jev-Omni chose Colander in the original order; reversing the choices produced Saucepan.

Across ten accepted object images, Jev-Omni matched **10/10 forward and 9/10 reversed**. This is useful as a prototype for explicit-choice visual quizzes or object sorting, while I06 shows why a plausible-looking result still needs checking. Accepted classifier JPEGs for these two examples are supplied for editing. The original corkscrew slot had been replaced by a desk fan before prediction after failed image acceptance; this was not a successful corkscrew test.

### Email: identify messages needing a human reply

The Japanese fictional email T01 asks the exhibition reception hours and whether entry without a reservation is possible. Its frozen reference is Reply required. T06 acknowledges receiving material and thanks the sender without an outstanding request; its reference is No reply required.

Jev-Omni matched **10/10 forward and 9/10 reversed** on ten fictional emails. T06 stayed correct in both orders, while T01 changed to No reply required under reversal. A possible use is proposing which messages belong in a human reply queue. No real inbox, automatic reply sending or proven labor saving was evaluated. The two original fictional email texts are supporting examples.

### Video: answer a question about a short visible sequence

The video example adds time to the same question-and-choice workflow. In V08, a lead car remains visibly distant and nonclosing; Jev-Omni chose Safe in both orders. V01 contains a visible immediate path conflict under the independently frozen Dangerous/Safe rubric, but Jev-Omni chose Safe in both orders. The question assessed any dangerous interval, including one followed by successful avoidance.

Jev-Omni matched **3/5 in each order** on five accepted synthetic clips. Both Dangerous references were missed. The result demonstrates that the video route produced decisions, and also gives a concrete failure worth showing. It does not justify delegating real driving or safety decisions to the model. In the main video, use V08 and a brief V01 counterexample; cars are one input example alongside the other three modalities.

## How these demonstrations were built

Scenarios, choice sets and reference labels were fixed before predictions. Fictional Japanese emails were authored directly; speech used Chatterbox Turbo, images used SDXL Lightning, and clips used H3. Generated inputs were checked independently before admission to the accepted set. Acceptance did not follow the model's answers.

Fixture generators were **ByteDance/SDXL-Lightning** (`sdxl_lightning_4step_unet.safetensors`, four steps, with the pinned SDXL base), **MiniMaxAI/MiniMax-H3** (Alibaba PDD8 runtime, recorded nine-step API setting), and **ResembleAI/chatterbox-turbo** (English default voice, no cloning). These generated the images, clips and speech inputs; **Jev-Omni and Clef were the decision evaluators**. Fictional emails were authored directly. Exact checkpoints/revisions, prompts, seeds and settings are recorded in `03-Technical-Appendices/GENERATOR_PROVENANCE.md` and `.json`; unexposed speech seeds remain unknown.

The study planned 40 inputs and accepted 35: ten speech, ten images, ten emails and five clips. Five planned clips remained unavailable. Jev-Omni completed 70 decisions, one original and one reversed choice order per accepted input. The primary forward result was **33/35 (94.3%)**; reversal was **31/35 (88.6%)**. These are accepted-subset observations. The second order repeats the same inputs and does not double the independent sample size.

| Jev-Omni task | Unique accepted inputs | Forward matches | Reversed matches |
| --- | ---: | ---: | ---: |
| Spoken complaint / inquiry | 10 | 10/10 | 10/10 |
| Object-image quiz | 10 | 10/10 | 9/10 |
| Japanese email reply decision | 10 | 10/10 | 9/10 |
| Short synthetic video decision | 5 | 3/5 | 3/5 |
| Total | 35 | 33/35 | 31/35 |

## Why compare with Clef?

The user wanted to compare the newly released Clef because its **Qwen/Qwen3.8-27B** backbone looked promising. That anticipated strength is the motivation for a test. The lineage is named by [Cloudflare's official Clef model card](https://huggingface.co/Cloudflare/clef). The experiment used the community quantization `simonlehmann/clef-NVFP4` at revision `817ac58ad358f42489980ead62f8bf7cafc628c2`; its [pinned model card](https://huggingface.co/simonlehmann/clef-NVFP4/blob/817ac58ad358f42489980ead62f8bf7cafc628c2/README.md) agrees. The downloaded pinned README/config matched the existing research hashes. A separate upstream Qwen revision was not recorded.

Clef was evaluated on the same 25 accepted image, text and video inputs, with the frozen questions and both choice orders. Jev-Omni's original saved responses were reused rather than rerun. **There was no Clef audio comparison.**

| Common task | Clef forward / reversed | Saved Jev-Omni forward / reversed |
| --- | ---: | ---: |
| Ten images | 10/10; 10/10 | 10/10; 9/10 |
| Ten emails | 10/10; 10/10 | 10/10; 9/10 |
| Five videos | 5/5; 5/5 | 3/5; 3/5 |
| Twenty-five inputs | 25/25; 25/25 | 23/25; 21/25 |

Across both orders, Clef matched **50/50 decisions**, versus **44/50 saved Jev-Omni decisions**, on 25 unique inputs. Clef had no label flips here; Jev-Omni had two, I06 and T01. Clef therefore performed better on this particular common cohort. Recommending Jev-Omni for its demonstrated four-modality workflow remains compatible with that result. Neither the Qwen backbone alone nor general model superiority is established by this comparison. The models received the same input bytes but used different video sampling and preprocessing.

## Recommendation and practical limits

Try Jev-Omni when the appeal is prototyping several input types around a common decision model. Start with explicit tasks, inspect actual inputs, and keep humans responsible for consequential decisions. For the image/text/video tasks covered here, Clef is also a comparison candidate worth testing; its measured result was stronger on the selected shared set. Evaluate an application's own examples before choosing either model.

This is a small, selected synthetic study with different tasks, languages and choice counts. It is not a general accuracy, safety, latency or cost benchmark. Choice order changed two Jev-Omni labels, and the video misses remain important limitations. Detailed collision-generation failures and emergency-stop amendments are optional technical appendices and do not supply additional main-story success claims. The closing takeaway is the usefulness of exploring one decision workflow across four input types, informed by a truthful comparison.

## Evidence references

- Jev-Omni results, frozen labels, acceptance corrections and runtime: `2026-10-02-jev-omni-40-synthetic/REPORT.md` and its `fixtures.json`, `accepted-selection.json`, `decisions.csv`.
- Clef common-cohort outcomes and preprocessing: `2026-10-02-clef-jev-comparison/REPORT.md` and its `decisions.csv`.
- Detailed speech caveats: `04-Audio-Sources/AUDIO_EXPERIMENT_MAPPING.md` in this delivered package.
- Exact source hashes and optional original-report extracts: `03-Technical-Appendices/RESEARCH_PROVENANCE_AND_LIMITS.md`.

Evidence reports are unchanged. No new inference, media generation or external publication accompanied this synthesis.
