Model specification · InclusionAI
Ling 3.0 Flash
inclusionai/ling-3.0-flash
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
SpecificationsReported values
- Context window
- 262K
- Maximum output
- 33K
- Architecture
- text->text
- Tokenizer
- Other
- Knowledge cutoff
- —
- Moderated
- No
- Input modalities
- text
- Output modalities
- text
Full pricingUSD where reported
| Price | Standard |
|---|---|
| Input | $0.021 / M |
| Output | $0.063 / M |
Benchmark indexesReported indexes
- Intelligence
- 20.6
- Coding
- 50.6
- Agentic
- 21
- Design Arena Elo
- —
Example workload costsStandard pricing only
| Workload | Input / output | Cache share | Estimated cost |
|---|---|---|---|
| Quick check | 10K / 1K × 1 | 0% | $0.00027 |
| Long analysis | 100K / 4K × 1 | 50% | $0.00151 |
| Workday | 20K / 2K × 1,000 | 30% | $0.445 |
More from InclusionAIAll InclusionAI models
| Model | Context | Input / M | Output / M | Intel. |
|---|---|---|---|---|
| Ling 3.0 Flash VL | 131K | $0.060 | $0.180 | 25 |
| Ling 3.0 Flash Fin | 262K | $0.060 | $0.180 | 23 |