Model library

Every frontier model, priced and ready

The complete Vidman AI lineup with its per-million-token rates and context windows — all callable through a single OpenAI-compatible endpoint. Swap the model id and ship; the rest of your code stays as it is.

Vidman AI Adaptive

vidman-adaptive

₹41.60 – ₹118.19/M Input · ₹124.81 – ₹373.47/M Output

1M Context

LLMToolsReasoningCached ₹13.24

Claude Opus 5

claude-opus-5

₹472.75/M Input · ₹2363.75/M Output

1M Context

LLMVisionToolsReasoningCached ₹47.28

Claude Sonnet 5

claude-sonnet-5

₹189.10/M Input · ₹945.50/M Output

1M Context

LLMVisionToolsReasoningCached ₹18.91

DeepSeek V4 Flash

deepseek-v4-flash

₹9.46/M Input · ₹23.64/M Output

128K Context

LLMToolsReasoningCached ₹0.95

DeepSeek V4 Pro

deepseek-v4-pro

₹132.37/M Input · ₹321.47/M Output

128K Context

LLMToolsReasoningCached ₹13.24
G

GLM 5.2

glm-5.2

₹132.37/M Input · ₹321.47/M Output

1M Context

LLMToolsReasoningCached ₹13.24

Kimi K2.7 Code

kimi-k2.7-code

₹75.64/M Input · ₹321.47/M Output

128K Context

LLMToolsReasoningCached ₹7.56

MiniMax M3

minimax-m3

₹28.37/M Input · ₹113.46/M Output

128K Context

LLMToolsReasoningCached ₹2.84

Qwen3.5 397B

qwen3.5-397b-a17b

₹47.28/M Input · ₹321.47/M Output

256K Context

LLMVisionToolsReasoningCached ₹4.73
D

Deepseek V4 Flash 0731

deepseek-v4-flash-0731

₹9.46/M Input · ₹23.64/M Output

128K Context

LLMToolsReasoningCached ₹0.95
D

Deepseek V4 Pro 0813

deepseek-v4-pro-0813

₹132.37/M Input · ₹321.47/M Output

128K Context

LLMToolsReasoningCached ₹13.24
G

Glm 5.3

glm-5.3

₹132.37/M Input · ₹416.02/M Output

1M Context

LLMToolsReasoningCached ₹24.58
G

Glm 5.3 Flash

glm-5.3-flash

₹14.18/M Input · ₹47.28/M Output

1M Context

LLMVisionToolsReasoningCached ₹2.84
K

Kimi K3

kimi-k3

₹283.65/M Input · ₹1418.25/M Output

1M Context

LLMVisionToolsReasoningCached ₹28.37
Q

Qwen3.8 2.4t A95b

qwen3.8-2.4t-a95b

₹189.10/M Input · ₹567.30/M Output

256K Context

LLMToolsReasoningCached ₹18.91

All figures in ₹ per 1M tokens, current as of 29 August 2026. Vidman AI Adaptive quotes a band because each request is billed at the model that served it; pinned models list their floor rate. Prompt caching bills separately. The dashboard shows the exact rate before a call goes out — that number wins over this page. ₹ at ₹94.55 per USD (committed rate, 2026-09-07 — live rate unavailable) — the same figure every ₹ above was converted at. Inference, fine-tuning, training, and dedicated hosting all run inside India.

On the roadmap

Text is only the first modality

Images and voice arrive on this same endpoint next. Create an account today and you will hear from us the day each one goes live.

Image generation

FComing soon

FLUX.2

Black Forest Labs

The open-weights quality bar for image generation, with fast klein variants for latency-sensitive work.

Billed per image · rates at launch

Coming soon

Qwen-Image-3.0

Alibaba

Text-heavy layouts -- posters, infographics, UI mockups -- from prompts up to 4.5K tokens.

Billed per image · rates at launch

SComing soon

Stable Diffusion 3.5 Large

Stability AI

The largest open ecosystem of fine-tunes, LoRAs and control tooling.

Billed per image · rates at launch

Text-to-speech

FComing soon

Fish Audio S2 Pro

Fish Audio

Open-weight quality leader with inline emotion control across ~50 languages.

Billed per 1M characters · rates at launch

Coming soon

Qwen3-TTS

Alibaba

Streaming speech in 10 languages, with voice cloning from a 3-second sample.

Billed per 1M characters · rates at launch

DComing soon

Dia2

Nari Labs

Streaming dialogue speech built for multi-speaker conversations.

Billed per 1M characters · rates at launch

Speech-to-text

Coming soon

Qwen3-ASR

Alibaba

Speech recognition across 52 languages, streaming and batch in one model.

Billed per audio minute · rates at launch

Coming soon

Voxtral Realtime

Mistral AI

Native streaming transcription at sub-second latency, Apache-2.0.

Billed per audio minute · rates at launch

WComing soon

Whisper large-v3 turbo

OpenAI

The most-deployed open transcription model, covering 99 languages.

Billed per audio minute · rates at launch

Voice-to-voice

Coming soon

PersonaPlex

NVIDIA

Full-duplex voice agents with persona and voice prompting; leads on task adherence.

Billed per audio minute · rates at launch

MComing soon

Moshi

Kyutai

The reference full-duplex speech model -- listens and speaks at once.

Billed per audio minute · rates at launch

Coming soon

Qwen3-Omni

Alibaba

One omni model taking text, image, audio and video in, streaming speech out.

Billed per audio minute · rates at launch

The lineup above is the open-weights release we are targeting per modality and can change before launch.