Sovereign AI for India
Secure. Sovereign. Sustainable.
Zero logs. Zero data retention.
Your prompts and responses exist in memory for the life of the request and are gone the moment it finishes. Nothing is written, stored or archived.
Prompts and responses live in memory for the duration of the request. No request logs, no content store, no archive.
We never train on your data, and there is no stored copy to breach, subpoena, or misuse. What was never kept cannot be lost.
We count tokens to bill you accurately. That count is all that exists afterwards -- never what you sent or what came back.
Enterprise
Compliance, control, and confidence — built in from day one, not bolted on.
Our engineers work alongside yours — designing the workload, integrating the systems, and staying until you are live. Not a handoff.
Vidman AI deploys where your data already lives: your cloud account, your data centre, or a hybrid of both. The endpoint does not change.
Zero retention of prompts and completions — they are never stored and never trained on. Separately, every administrative action carries a full audit trail, so your risk team can see who did what.
Sovereign by design
Vidman AI is operated by VTT AI Private Limited, Bangalore. The serving stack is Indian technology, and this site is served from Mumbai.
Every request you send, every fine-tuning and training run, and every model we host for you executes inside the country — on Indian servers, operated by an Indian company. Not a region you selected. The whole stack.
Prompts and responses live in memory for the life of a request and are gone when it completes. Nothing stored, nothing trained on — in India or anywhere else.
What it costs
Save ₹1,30,00,625/mo
on deepseek-v4-pro vs Nebius
12%
vs Together AI
12%
vs Nebius
Pick the model you use most
Choose a period, then a volume — or type your own
What share of your tokens are input rather than output
Based on 500 billion tokens/mo · 350 billion input, 150 billion output
| Input / 1M | Output / 1M | Monthly | |
|---|---|---|---|
| Vidman AI | ₹132.37 | ₹321.47 | ₹9,45,50,000 |
| Together AI | ₹164.52 | ₹329.03 | ₹10,69,36,050 |
| Nebius | ₹165.46 | ₹330.93 | ₹10,75,50,625 |
Together AI serverless, as published 29 August 2026.
Split on which model to pick? let Adaptive make that call for you
Long-context models billed per million tokens. Swap one string — the model id — and the rest of your code carries on unchanged.
| Model | Context | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| vidman-adaptive | 1M | ₹41.60 – ₹118.19 | ₹13.24 | ₹124.81 – ₹373.47 |
| claude-opus-5 | 1M | ₹472.75 | ₹47.28 | ₹2363.75 |
| claude-sonnet-5 | 1M | ₹189.10 | ₹18.91 | ₹945.50 |
| deepseek-v4-flash | 128K | ₹9.46 | ₹0.95 | ₹23.64 |
| deepseek-v4-flash-0731 | 128K | ₹9.46 | ₹0.95 | ₹23.64 |
| deepseek-v4-pro | 128K | ₹132.37 | ₹13.24 | ₹321.47 |
| deepseek-v4-pro-0813 | 128K | ₹132.37 | ₹13.24 | ₹321.47 |
| glm-5.2 | 1M | ₹132.37 | ₹13.24 | ₹321.47 |
| glm-5.3 | 1M | ₹132.37 | ₹24.58 | ₹416.02 |
| glm-5.3-flash | 1M | ₹14.18 | ₹2.84 | ₹47.28 |
| kimi-k2.7-code | 128K | ₹75.64 | ₹7.56 | ₹321.47 |
| kimi-k3 | 1M | ₹283.65 | ₹28.37 | ₹1418.25 |
| minimax-m3 | 128K | ₹28.37 | ₹2.84 | ₹113.46 |
| qwen3.5-397b-a17b | 256K | ₹47.28 | ₹4.73 | ₹321.47 |
| qwen3.8-2.4t-a95b | 256K | ₹189.10 | ₹18.91 | ₹567.30 |
₹ per 1M tokens, as of 29 August 2026. Vidman AI Adaptive shows a range because its rate follows the model each request lands on; pinned models show their floor. Prompt caching is billed separately. The exact rate for the model you are about to call is shown in the dashboard before you send a request — treat that as authoritative over this table.
Agent integrations
The frameworks your team ships with today work as-is — one OpenAI-compatible endpoint, and no new SDK to pick up.
Custom models
Fine-tuning here ends in a file you hold. That checkpoint can move straight onto its own dedicated endpoint without leaving the platform.
Reached through the same OpenAI-compatible API as everything else. This is part of fine-tuning rather than the serverless product — different billing, and a different reason to reach for it.
A LoRA or QLoRA adapter is served with its base model automatically, with no manual merge step.
Chat, completion, embedding or reranker — each exposes the matching OpenAI route.
Light fits on the minimum that will hold the model; heavy adds GPUs for concurrency.
bf16 or fp8, with KV cache compression to fit more concurrent requests on the same card.
Vidman AI Adaptive
Adaptive hands each request to the model that fits it — quality, speed and budget weighed per call — on per-token rates with the ceiling printed up front.
Which one you need
A fine-tune costs more than a serverless call and takes longer to get right — and plenty of workloads belong on serverless for good. Two cards to place yourself.
A general-purpose model is already handling it — or the job itself is still taking shape.
A general-purpose model gets close and misses in a pattern you can name.
Enterprise
The whole stack — OpenAI-compatible endpoint, the model catalog, adaptive routing — on your infrastructure, on-prem or in your cloud. The work never crosses your boundary; we run the machinery.
Put a production model to work over the serverless endpoint, or train your own on dedicated GPUs. Two products, one account, nothing committed on either side.
Serverless by the million tokens · Fine-tuning from ₹$39.71/hr · Your weights stay yours