Sovereign AI for India

India runs on Vidman AI

Secure. Sovereign. Sustainable.

Zero logs. Zero data retention.

The models worth building on

Nothing about your data is kept

Your prompts and responses exist in memory for the life of the request and are gone the moment it finishes. Nothing is written, stored or archived.

Nothing Written Down

Prompts and responses live in memory for the duration of the request. No request logs, no content store, no archive.

Nothing to Leak

We never train on your data, and there is no stored copy to breach, subpoena, or misuse. What was never kept cannot be lost.

Metering, Not Logging

We count tokens to bill you accurately. That count is all that exists afterwards -- never what you sent or what came back.

Enterprise

Enterprise-grade. Out of the box.

Compliance, control, and confidence — built in from day one, not bolted on.

Forward deployed

Our engineers work alongside yours — designing the workload, integrating the systems, and staying until you are live. Not a handoff.

  • A dedicated engineer from day one
  • Joint design and build
  • Ongoing accuracy and cost tuning
  • Production support, SLA-backed

Deployment flexibility

Vidman AI deploys where your data already lives: your cloud account, your data centre, or a hybrid of both. The endpoint does not change.

  • Private cloud, on-premise, or hybrid
  • Air-gapped deployment available
  • Bring your own model, or use the catalog
  • One OpenAI-compatible API — swap without rewriting
  • Fine-tuned weights stay yours, and exportable

Security and governance

Zero retention of prompts and completions — they are never stored and never trained on. Separately, every administrative action carries a full audit trail, so your risk team can see who did what.

  • Zero logs, zero data retention
  • Role-based access control
  • Audit trail on administrative actions
  • SOC 2 Type II, ISO 27001, DPDP compliant
  • SOC 2 Type II
  • ISO 27001
  • DPDP compliant
  • Role-based access
  • Audit trail
  • Built and hosted in India

Sovereign by design

An Indian platform that says where it runs

An Indian company

Vidman AI is operated by VTT AI Private Limited, Bangalore. The serving stack is Indian technology, and this site is served from Mumbai.

Every workload runs in India

Every request you send, every fine-tuning and training run, and every model we host for you executes inside the country — on Indian servers, operated by an Indian company. Not a region you selected. The whole stack.

Nothing retained

Prompts and responses live in memory for the life of a request and are gone when it completes. Nothing stored, nothing trained on — in India or anywhere else.

What does “sovereign” actually mean here?

Where does inference run?
Inside India. Every call you make is served from Indian infrastructure, whichever model answers it, and every fine-tuning and training run executes here too.
Who operates the stack?
Vidman AI — an Indian company. On managed capacity, or inside your own environment on an enterprise deployment.
What is retained?
Nothing. Prompts and responses are processed in memory and discarded when the request completes.

What it costs

Run the numbers before you commit a token

Save ₹1,30,00,625/mo

on deepseek-v4-pro vs Nebius

12%

vs Together AI

12%

vs Nebius

Pick the model you use most

Choose a period, then a volume — or type your own

tokens/mo

What share of your tokens are input rather than output

Input 70%Output 30%

Based on 500 billion tokens/mo · 350 billion input, 150 billion output

Input / 1MOutput / 1MMonthly
Vidman AI₹132.37₹321.47₹9,45,50,000
Together AI₹164.52₹329.03₹10,69,36,050
Nebius₹165.46₹330.93₹10,75,50,625

Together AI serverless, as published 29 August 2026.

Get 50% extra on your first wallet top-up.

Billed per second of GPU time.

Split on which model to pick? let Adaptive make that call for you

The catalog, with its prices on the sleeve

Long-context models billed per million tokens. Swap one string — the model id — and the rest of your code carries on unchanged.

Six families
Claude · DeepSeek · GLM · Kimi · MiniMax · Qwen
1M tokens
Longest context window
₹9.46
Lowest input, per 1M tokens
Text + vision
Models that accept image input

₹ per 1M tokens, as of 29 August 2026. Vidman AI Adaptive shows a range because its rate follows the model each request lands on; pinned models show their floor. Prompt caching is billed separately. The exact rate for the model you are about to call is shown in the dashboard before you send a request — treat that as authoritative over this table.

Agent integrations

Your agent stack already speaks to us

The frameworks your team ships with today work as-is — one OpenAI-compatible endpoint, and no new SDK to pick up.

Hermes AgentOpenClawClineLiteLLMOpenAI SDKAnthropic SDK

Custom models

Your weights, on a GPU of their own

Fine-tuning here ends in a file you hold. That checkpoint can move straight onto its own dedicated endpoint without leaving the platform.

Custom training, end to end

Comes with fine-tuning

Custom Model Endpoints

Reached through the same OpenAI-compatible API as everything else. This is part of fine-tuning rather than the serverless product — different billing, and a different reason to reach for it.

Per second of GPU time
The same rates as training — from ₹39.71/hr across 12 GPU types

Adapters serve themselves

A LoRA or QLoRA adapter is served with its base model automatically, with no manual merge step.

The route matches the model

Chat, completion, embedding or reranker — each exposes the matching OpenAI route.

Sized to expected traffic

Light fits on the minimum that will hold the model; heavy adds GPUs for concurrency.

Memory is yours to tune

bf16 or fp8, with KV cache compression to fit more concurrent requests on the same card.

Vidman AI Adaptive

One endpoint. Every model. None of the model-ops.

Adaptive hands each request to the model that fits it — quality, speed and budget weighed per call — on per-token rates with the ceiling printed up front.

Which one you need

Should you fine-tune at all?

A fine-tune costs more than a serverless call and takes longer to get right — and plenty of workloads belong on serverless for good. Two cards to place yourself.

Stay on serverless inference

A general-purpose model is already handling it — or the job itself is still taking shape.

  • You are still validating the idea or shipping v1
  • The prompts themselves still shift weekly
  • Traffic spikes, follows a season, or stays small
  • Something in the catalog already meets the accuracy bar
  • You hold no labelled examples of the output you want

Fine-tune your own model

A general-purpose model gets close and misses in a pattern you can name.

  • You own examples of the desired output — hundreds or thousands, not millions
  • A smaller, cheaper model could own one narrow job
  • The base model has never met your domain language — clinical, legal, in-house jargon
  • Format and tone must land on every call
  • The weights must be yours

Enterprise

Want it inside your own walls?

The whole stack — OpenAI-compatible endpoint, the model catalog, adaptive routing — on your infrastructure, on-prem or in your cloud. The work never crosses your boundary; we run the machinery.

Begin with either one

Put a production model to work over the serverless endpoint, or train your own on dedicated GPUs. Two products, one account, nothing committed on either side.

Serverless by the million tokens · Fine-tuning from ₹$39.71/hr · Your weights stay yours