Llama 3.3 70B Instruct
llama-3.3-70b-instruct
Open-weights flagship and the sensible default. Strong instruction following across eight languages with tool calling.
- provider
- Meta
- context
- 128k
- speed
- balanced
catalog
These ids map relay requests to a configured OpenAI-compatible provider. Catalog inclusion is not a claim of availability; effective status is shown on every card and by GET /v1/models.
llama-3.3-70b-instruct
Open-weights flagship and the sensible default. Strong instruction following across eight languages with tool calling.
llama-3.1-8b-instruct
The workhorse small model. Cheap, quick and reliable for classification, extraction and routing.
qwen-2.5-72b-instruct
Excellent multilingual coverage and structured output. A strong pick for agentic workflows.
qwen-2.5-coder-32b
Repository-scale code generation, diff editing and test writing. Currently the best open code model per parameter.
mistral-small-3.1
Efficient multimodal model with vision support and a very low cost per token. Good latency-to-quality ratio.
mistral-nemo
Compact long-context model built with NVIDIA. Handles 128k context comfortably for summarisation tasks.
deepseek-r1-distill-70b
Reasoning distilled from R1 into Llama 70B. Emits an explicit chain of thought before answering.
gemma-3-27b-it
Compact multimodal model tuned for helpful, grounded responses with image understanding.
gemma-2-9b-it
Very small footprint, surprisingly good writing quality. Good for high-volume, low-stakes text.
phi-4
Synthetic-data trained reasoning model that punches well above its size on maths and logic.