OmniJev

视觉输入,毫秒级决策输出Visual input. Millisecond decisions.

OmniJev 是我们在自建数据集上端到端训练的 4B 决策模型。
它联合读取文本、图像与视频,在一次前向计算中为 choice、score、noul 问题输出校准概率。
OmniJev is a 4B decision model trained end to end on our own dataset.
It reads text, images and video together, returning calibrated probabilities for choice, score and noul questions in one forward pass.

0生成 tokengenerated tokens103 ms单问题时延中位数median single-question latency3模态3 modalities图像视频文本image, video, text

模型特性Model properties

训练得到的决策模型A trained decision model

输入由视觉状态、文本上下文和问题集合组成。模型直接输出固定结构的概率结果,不调用第三方模型,也不进行逐 token 文本生成。Each request contains visual state, text context and typed questions. The model itself returns fixed probability structures—without third-party model calls or token-by-token generation.

类型化答案Typed answers

支持选择、评分与真伪判断,结果天然符合结构,无需解析生成文本。Choice, score and boolean answers that fit your schema by construction.

概率经过校准Calibrated confidence

回答 0.9 时,大约十次对九次。阈值因此真正具备业务意义。When it says 0.9, it is right roughly nine times out of ten.

毫秒级响应Milliseconds

一次前向同时回答一组问题,图像单题服务端时延中位数仅 103 毫秒。One forward pass answers a full set of questions; median server latency for one image question is 103 ms.

全模态输入Omni-modal input

统一处理照片、连续视频、应用界面、网页、棋盘与具身视角。Photos, videos, interfaces, web pages, games and embodied views.

可控地拒答Built-in abstention

每个选择都包含“以上都不是”,不确定的样本可安全交给其他流程。Every choice can return “none of the above” for safe routing.

端到端训练End-to-end trained

约 27 万条决策记录、130 万个类型化问题,覆盖网页与手机、机器人、视频、游戏、手势、危险和声音。About 270k decision records and 1.3M typed questions spanning web and phone use, robots, video, games, gestures, hazards and sounds.

视觉输入Visual input使用场景Use cases
界面、应用、网站Interfaces, apps, websites下一步动作、目标元素、操作的不可逆程度、当前是否出错Next action, target element, irreversibility and error state
视频Video当前行为、事件时间、下一步事件、指定事件是否发生Activity, timing, next event and whether an event occurred
通用场景Scenes方位、尺寸、距离、数量,以及指定描述是否成立Direction, size, distance, count and statement validity
标注区域Marked regions符合描述的区域及其置信度Which region matches a description and with what confidence
具身视角Embodied views当前执行的指令、下一步移动方向、是否应立即操作Current instruction, next direction and whether to act now
游戏与棋盘Games and boards下一步动作、合法性、威胁和棋子位置Next move, legality, threats and piece location
安全与手势Safety and gestures火、烟或武器是否出现,以及手势类别和对应指令Fire, smoke, weapons, gesture class and mapped command

任务示例Task examples

图像与视频决策任务Image and video decision tasks

以下视频展示机器人任务、四类实时游戏、网页与手机操作、目标定位、视频理解和安全监测。The examples cover robot tasks, four real-time games, web and phone operation, visual grounding, video understanding and safety monitoring.

ROBOTICS

四个机器人任务,同时判断Four robot tasks at once

每一帧同步判断子任务、运动方向、抓取状态与阶段完成度。Every frame resolves sub-task, motion, grasp state and completion.

ROBOT DECISIONS

长程机器人任务,逐帧详解A long-horizon robot task, frame by frame

以 LIBERO-10 为例,逐帧展示任务与子任务识别、运动控制,以及状态判断。A LIBERO-10 example showing task and sub-task recognition, seven-way action probabilities, grasp state, object holding and completion on every frame.

REAL-TIME GAMES

四种游戏,逐帧决策Four games, frame by frame

在 Atari、贪吃蛇、五子棋和国际象棋中判断下一步动作、得分机会与局势。Decide the next action, scoring opportunity and game state in Atari, Snake, Gomoku and Chess.

WEB AGENT

21 步网页任务A 21-step web task

逐页选择操作元素、动作类型、风险与任务进度。Pick the element, operation, risk and progress on every page.

PHONE AGENT

手机任务,逐屏操作Operating a phone, screen by screen

逐屏判断下一步动作、任务是否完成、操作风险与当前进度。Choose the next action, completion state, operation risk and progress on each screen.

GROUNDING

无需手工框选的指点Pointing without boxes

用网格在截图或照片中定位任意自然语言所指的对象。Locate a natural-language target on screenshots or photos.

VIDEO

视频事件理解Video event understanding

判断灯是否点亮、门是否开合,以及人物最后的状态。Understand lights, doors, people and temporal events.

SAFETY

监控画面中的危险Hazards in a camera feed

识别火、烟、危险类型、紧急程度与所在区域。Fire, smoke, urgency and the affected region.

Held-out evaluation

基准测试结果Benchmark results

展示五组模型结果均完整的 held-out 评测项:比较 Qwen3.5 zero-shot 骨干与三个规模的 OmniJev。准确率按百分比展示;ECE 为 OmniJev-4B 的校准误差,越低越好。Held-out benchmarks with results for all five models: Qwen3.5 zero-shot backbones and three OmniJev sizes. Accuracy is shown as a percentage; ECE is the calibration error of OmniJev-4B and lower is better.

Qwen3.5-0.8B zero-shotOmniJev-0.8BOmniJev-2BQwen3.5-4B zero-shotOmniJev-4Baccuracy ↑  ·  ECE ↓
LIBERO-10机器人决策,每帧 8 题 · n=1,504Robot decisions, 8 questions per frame · n=1,504
55.1
Qwen3.5-0.8B ZS
77.1
OmniJev-0.8B
72.4
OmniJev-2B
29.9
Qwen3.5-4B ZS
80.7
OmniJev-4B
ECE0.031
Mind2Web testtask / website / domain · n=1,500
40.1
Qwen3.5-0.8B ZS
59.6
OmniJev-0.8B
63.1
OmniJev-2B
30.3
Qwen3.5-4B ZS
73.3
OmniJev-4B
ECE0.038
Grid pointing网页 96 格目标定位 · n=1,50096-cell web grounding · n=1,500
38.4
Qwen3.5-0.8B ZS
47.4
OmniJev-0.8B
62.5
OmniJev-2B
42.1
Qwen3.5-4B ZS
73.7
OmniJev-4B
ECE0.021
Charades-STA视频事件定位 · n=1,500Video event localization · n=1,500
39.0
Qwen3.5-0.8B ZS
81.1
OmniJev-0.8B
82.6
OmniJev-2B
54.3
Qwen3.5-4B ZS
85.9
OmniJev-4B
ECE0.021
Catch game逐帧游戏决策 · n=1,504Frame-level game decisions · n=1,504
40.9
Qwen3.5-0.8B ZS
66.1
OmniJev-0.8B
72.8
OmniJev-2B
15.1
Qwen3.5-4B ZS
87.0
OmniJev-4B
ECE0.016
HaGRID + safety手势、火、烟与武器 · n=1,500Gestures, fire, smoke and weapons · n=1,500
52.9
Qwen3.5-0.8B ZS
96.9
OmniJev-0.8B
98.4
OmniJev-2B
68.3
Qwen3.5-4B ZS
98.7
OmniJev-4B
ECE0.009
OK-VQA答案候选池 · n=1,500Answer pool · n=1,500
79.3
Qwen3.5-0.8B ZS
65.6
OmniJev-0.8B
76.5
OmniJev-2B
86.0
Qwen3.5-4B ZS
80.9
OmniJev-4B
ECE0.021
LongVideoBench val长视频理解 · n=500Long-video understanding · n=500
41.0
Qwen3.5-0.8B ZS
48.6
OmniJev-0.8B
50.2
OmniJev-2B
58.5
Qwen3.5-4B ZS
58.2
OmniJev-4B
ECE0.073
Long video / planning / spatial长视频、规划与空间 · n=1,511Long video, planning and spatial · n=1,511
39.0
Qwen3.5-0.8B ZS
49.4
OmniJev-0.8B
52.9
OmniJev-2B
48.6
Qwen3.5-4B ZS
60.2
OmniJev-4B
ECE0.072
Mixed decision set区域、存在性、手机、棋类与短视频 · n=1,491Regions, existence, phone, games and short video · n=1,491
54.5
Qwen3.5-0.8B ZS
60.4
OmniJev-0.8B
62.3
OmniJev-2B
58.2
Qwen3.5-4B ZS
66.1
OmniJev-4B
ECE0.079
Wiki navigation网页导航 · n=1,500Web navigation · n=1,500
41.4
Qwen3.5-0.8B ZS
65.9
OmniJev-0.8B
66.7
OmniJev-2B
33.6
Qwen3.5-4B ZS
70.6
OmniJev-4B
ECE0.030
Real-robot MUTEX真实机器人任务 · n=1,500Real-robot tasks · n=1,500
44.2
Qwen3.5-0.8B ZS
68.1
OmniJev-0.8B
68.7
OmniJev-2B
26.7
Qwen3.5-4B ZS
74.9
OmniJev-4B
ECE0.022

注:此处仅展示当前所有对比列均完整的评测项,缺项 benchmark 暂不纳入。结果会随评测与复核持续更新,并沿用一致协议以确保可比;完整结果见 GitHub Results ↗。Note: Only benchmarks with complete comparison columns are shown; incomplete rows are omitted for now. Results will be updated under consistent protocols as evaluation and verification continue. See the full GitHub results ↗.

推理延迟Inference latency

服务端实测数据Measured server latency

测试模型:OmniJev-4BModel tested: OmniJev-4B时延均报告 12 次运行的中位数。Latency values are medians over 12 runs.

01103ms

单图单题,服务端时延one image, one question

02153ms

同图十二题,总时延twelve questions, total latency

031s

整段视频,多帧采样a full sampled video

040token

始终不生成文本generated, always

Online evaluation

输入图片/视频进行体验Evaluate an image or video

在线页面支持 choice、score、noul 问题,并显示完整概率分布和服务端耗时。The online page supports choice, score and noul questions and reports full probabilities and server latency.