Model specification · Xiaomi
MiMo-V2.5
xiaomi/mimo-v2.5
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
SpecificationsReported values
- Context window
- 1.1M
- Maximum output
- 131K
- Architecture
- text+image+audio+video->text
- Tokenizer
- Other
- Knowledge cutoff
- —
- Moderated
- No
- Input modalities
- text audio image video
- Output modalities
- text
Full pricingUSD where reported
| Price | Standard |
|---|---|
| Input | $0.140 / M |
| Output | $0.280 / M |
Benchmark indexesReported indexes
- Intelligence
- 37.2
- Coding
- 56.8
- Agentic
- 23.7
- Design Arena Elo
- 1265
Example workload costsStandard pricing only
| Workload | Input / output | Cache share | Estimated cost |
|---|---|---|---|
| Quick check | 10K / 1K × 1 | 0% | $0.00168 |
| Long analysis | 100K / 4K × 1 | 50% | $0.00826 |
| Workday | 20K / 2K × 1,000 | 30% | $2.54 |
More from XiaomiAll Xiaomi models
| Model | Context | Input / M | Output / M | Intel. |
|---|---|---|---|---|
| MiMo-V2.5-Pro | 1.1M | $0.435 | $0.870 | 42.2 |