Model library
Every frontier model, priced and ready
The complete Vidman AI lineup with its per-million-token rates and context windows — all callable through a single OpenAI-compatible endpoint. Swap the model id and ship; the rest of your code stays as it is.
Vidman AI Adaptive
vidman-adaptive
₹41.60 – ₹118.19/M Input · ₹124.81 – ₹373.47/M Output
1M Context
Claude Opus 5
claude-opus-5
₹472.75/M Input · ₹2363.75/M Output
1M Context
Claude Sonnet 5
claude-sonnet-5
₹189.10/M Input · ₹945.50/M Output
1M Context
DeepSeek V4 Flash
deepseek-v4-flash
₹9.46/M Input · ₹23.64/M Output
128K Context
DeepSeek V4 Pro
deepseek-v4-pro
₹132.37/M Input · ₹321.47/M Output
128K Context
GLM 5.2
glm-5.2
₹132.37/M Input · ₹321.47/M Output
1M Context
Kimi K2.7 Code
kimi-k2.7-code
₹75.64/M Input · ₹321.47/M Output
128K Context
MiniMax M3
minimax-m3
₹28.37/M Input · ₹113.46/M Output
128K Context
Qwen3.5 397B
qwen3.5-397b-a17b
₹47.28/M Input · ₹321.47/M Output
256K Context
Deepseek V4 Flash 0731
deepseek-v4-flash-0731
₹9.46/M Input · ₹23.64/M Output
128K Context
Deepseek V4 Pro 0813
deepseek-v4-pro-0813
₹132.37/M Input · ₹321.47/M Output
128K Context
Glm 5.3
glm-5.3
₹132.37/M Input · ₹416.02/M Output
1M Context
Glm 5.3 Flash
glm-5.3-flash
₹14.18/M Input · ₹47.28/M Output
1M Context
Kimi K3
kimi-k3
₹283.65/M Input · ₹1418.25/M Output
1M Context
Qwen3.8 2.4t A95b
qwen3.8-2.4t-a95b
₹189.10/M Input · ₹567.30/M Output
256K Context
All figures in ₹ per 1M tokens, current as of 29 August 2026. Vidman AI Adaptive quotes a band because each request is billed at the model that served it; pinned models list their floor rate. Prompt caching bills separately. The dashboard shows the exact rate before a call goes out — that number wins over this page. ₹ at ₹94.55 per USD (committed rate, 2026-09-07 — live rate unavailable) — the same figure every ₹ above was converted at. Inference, fine-tuning, training, and dedicated hosting all run inside India.
On the roadmap
Text is only the first modality
Images and voice arrive on this same endpoint next. Create an account today and you will hear from us the day each one goes live.
Image generation
FLUX.2
Black Forest Labs
The open-weights quality bar for image generation, with fast klein variants for latency-sensitive work.
Billed per image · rates at launch
Qwen-Image-3.0
Alibaba
Text-heavy layouts -- posters, infographics, UI mockups -- from prompts up to 4.5K tokens.
Billed per image · rates at launch
Stable Diffusion 3.5 Large
Stability AI
The largest open ecosystem of fine-tunes, LoRAs and control tooling.
Billed per image · rates at launch
Text-to-speech
Fish Audio S2 Pro
Fish Audio
Open-weight quality leader with inline emotion control across ~50 languages.
Billed per 1M characters · rates at launch
Qwen3-TTS
Alibaba
Streaming speech in 10 languages, with voice cloning from a 3-second sample.
Billed per 1M characters · rates at launch
Dia2
Nari Labs
Streaming dialogue speech built for multi-speaker conversations.
Billed per 1M characters · rates at launch
Speech-to-text
Qwen3-ASR
Alibaba
Speech recognition across 52 languages, streaming and batch in one model.
Billed per audio minute · rates at launch
Voxtral Realtime
Mistral AI
Native streaming transcription at sub-second latency, Apache-2.0.
Billed per audio minute · rates at launch
Whisper large-v3 turbo
OpenAI
The most-deployed open transcription model, covering 99 languages.
Billed per audio minute · rates at launch
Voice-to-voice
PersonaPlex
NVIDIA
Full-duplex voice agents with persona and voice prompting; leads on task adherence.
Billed per audio minute · rates at launch
Moshi
Kyutai
The reference full-duplex speech model -- listens and speaks at once.
Billed per audio minute · rates at launch
Qwen3-Omni
Alibaba
One omni model taking text, image, audio and video in, streaming speech out.
Billed per audio minute · rates at launch
The lineup above is the open-weights release we are targeting per modality and can change before launch.