Built for builders

Build Without

Limits

Access the best AI models, run them reliably, and scale globally — all through one powerful API.

The MiniCrow mascot

Models
take you further.

  • One APIText, speech, video, retrieval
  • Routing you can readEvery call says why
  • Priced per callCost in the response
  • PrepaidZero balance, hard stop

Language models

Three tiers. Pick the one that fits the job.

What you buy is a lane, and the lane is what stays stable while the model behind it is repriced, re-quantised or replaced. Every response still tells you which lane answered, the rule that picked it, and what the call cost — a router that hides its choice cannot be debugged.

Flash Lite

osprey-flash-lite

Not available

The cheap one that still sees and hears.

One model, two speeds: thinking off for volume, thinking on when the answer has to be right. It takes text, images, audio and files, which for its price is the surprise of the catalogue — but it will not take tools.

Good for

  • High-volume classification
  • Tagging and routing
  • Cheap extraction passes
  • Describing an attachment

Modes

  • :speed₹9.5 · ₹35.9

    Thinking off. For volume.

    Takes text · images · audio · files

  • :intelligencedefault₹9.5 · ₹35.9

    Thinking on. The default.

    Takes text · images · audio · files

No function calling. Tools are refused with a 400, not silently dropped. It is an owner's decision for v1 and it is stated rather than discovered halfway through an agent run.

₹ per million tokens, in · out

Priced and in the catalogue. The self-hosted endpoint behind it is not deployed yet.

Flash

osprey-flash

Live

The workhorse. Nearly everything belongs here.

Three lanes across a 3× price range, with tools on every one of them. This is the tier the router was designed around: fast and cheap for short exchanges, the middle lane the moment tools appear, and a premium rung you can name when an answer has to hold up.

Good for

  • Chat and assistants
  • Agent loops with tools
  • Structured extraction
  • Reading screenshots
  • Summarising a long document

Modes

  • :speed₹6.9 · ₹19.0

    Short exchanges, low latency.

    Takes text only

  • :intelligencedefault₹7.9 · ₹26.4

    The default. Tools, images, most work.

    Takes text · images

  • :max₹21.1 · ₹126.7

    When it has to hold up.

    Takes text · images · files

Function calling on every mode.

₹ per million tokens, in · out

Pro

osprey-pro

Live

A different price class, not a nicer Flash.

Reach for Pro when the work is a long agent loop with real consequences, or when a flash-class model has already been tried and failed. Its middle lane costs roughly nineteen times Flash's, which is the honest way to describe it: this is not an upgrade, it is a different budget.

Good for

  • Multi-turn agent loops
  • Long-context reasoning
  • Code and analysis that must be right
  • Audio and video in one prompt

Modes

  • :speed₹79.2 · ₹396.0

    The widest input of any lane.

    Takes text · images · audio · files

  • :intelligencedefault₹147.8 · ₹464.6

    The default. Deep reasoning.

    Takes text only

  • :max₹211.2 · ₹1,056.0

    The top of the catalogue.

    Takes text · images · files

Function calling on every mode.

₹ per million tokens, in · out

Every language lane, side by side

1M+ context, every lane
LaneBest forTakesTools₹/Mtok in₹/Mtok out
osprey-flash-liteVolume classificationtext · image · audio · file9.535.9
osprey-flash:speedShort chat, low latencytext6.919.0
osprey-flash:intelligenceTools, images, most worktext · image7.926.4
osprey-flash:maxAnswers that must hold uptext · image · file21.1126.7
osprey-pro:speedWidest input, fasttext · image · audio · file79.2396.0
osprey-pro:intelligenceDeep reasoningtext147.8464.6
osprey-pro:maxLong agent loopstext · image · file211.21,056.0
Third-party figure

Flash-class models collapse on multi-turn tool use.

On BFCL v4's multi-turn split, flash-class models score 13–36% where pro-class models score 60–68%. It is not a uniform five-point tax — it is a cliff, and it is why a wrong cheap route breaks an agent loop while a wrong expensive one only costs markup.

Berkeley Function-Calling Leaderboard v4 — published third-party, not our measurement

No board of our own

We do not publish a benchmark board of our own, because we do not have one yet and borrowing someone else's would be worse than saying so. No public board puts these lanes side by side, and not one of them reports a Hinglish or Indic figure — which is most of the traffic this gateway was built for. Every call's route reason and cost is logged, and that is what our own numbers will be built from.

Speech, video and retrieval

The rest of the catalogue, built on Indian audio.

Transcription, video summaries, speech and vector search — the same key, the same prepaid balance, the same cost in the response.

Live

Lark

Speech to text

Indian-language transcription, code-mix included. Measured on real 8 kHz call audio, not on studio recordings.

Good for

  • Phone-call transcription
  • Voice notes and meetings
  • Hinglish and Indic speech
  • Support-call QA
  • Anything recorded at 8 kHz

POST /v1/audio/transcriptions

  • lark-nanoNot availableSelf-hosted₹7.07per hour of audio
  • lark-miniDefault lane₹7.07per hour of audio
  • lark-largeWider output budget₹42.50per hour of audio
Live

Lark-V

Video

A summary and a seekable timeline from one upload. The large tier watches the clip natively — it sees the frames and hears the speech, in the script it was spoken in.

Good for

  • Video summaries
  • A seekable timeline of a clip
  • Screen recordings and demos
  • Ad and creative review
  • Spoken content inside video

POST /v1/video/summaries

  • lark-v-nanoNot availableTiled framesvision + speech+20% on the sum
  • lark-v-miniNot availableTiled framesvision + speech+20% on the sum
  • lark-v-largeNative video, sees and hearsupstream + 20%measured per clip

The large tier is live. The two frame-tiling tiers are not deployed and say so rather than upgrading you to it.

Live

Pica

Text to speech

Eleven ready voices — seven Hindi, four English — each a sha256-pinned reference. Two lanes: standard for both languages, expressive for Hindi with emotion. Two delivery modes on each.

Good for

  • Voice notifications and IVR
  • Hindi narration with emotion
  • Audio versions of written content
  • Product and demo voiceover
  • Accessibility read-aloud

POST /v1/audio/speech

  • pica-ministandard and expressive lanes₹36per 10,000 characters
  • pica-largeNot availableAnnounced, not wired up yetprice not set

Embeddings and rerank

Live

Dense and sparse vectors from one call, and a cross-encoder that reranks a shortlist. Unbranded on purpose — this one has not been given a name yet.

  • Search over your own documents
  • RAG retrieval
  • Deduplication
  • Reranking a shortlist

POST /v1/embeddings · POST /v1/rerank

₹2.1

per Mtok · dense + sparse

Prices are the catalogue's seed rates, at a 20% markup over what a call costs us and ₹88 to the dollar — a configured rate, not a live FX feed. The operator panel changes any of them, and a change never rewrites a call already made.

Auto mode

The router tells you what it picked, and why.

It is rules, not a model — structural signals read off the request body in under two milliseconds, with no English keyword lists, because a keyword list scores Hinglish, Gujarati and Tamil traffic identically to noise. Every decision comes back with the same reason id it was recorded under, so a route you disagree with is a string you can search for rather than a mood you have to argue with.

H6:open_tool_loop

Never switch mode inside an open tool loop

Thought signatures and thinking blocks bind a continuation to the model that began it. All three upstreams 400 when replayed elsewhere.

S:tools

Tools present, or three turns deep, takes the middle lane

A wrong cheap route breaks the agent loop. A wrong expensive one only costs markup.

S:sticky

A thread keeps the mode it started on

A switch throws away the prompt cache, and cache-read is a fraction of input price on every lane.

S:long_input

Eight thousand tokens of input is a summarisation job

One long user turn with no tools is a different shape of work from a conversation.

H2

A lane that cannot take the modality is removed

An image on a text-only lane is not a worse answer, it is an error.

E

One rung up, once, only on a verifiable failure

A tool call that cannot be executed is evidence. A truncated answer is not — that is a continuation problem in the same mode.

200 · POST /v1/chat/completions
{
  "model": "osprey-flash",
  "x_minicrow": {
    "requested_mode": "auto",
    "served_mode":    "intelligence",
    "route_reason":   "S:tools",
    "cost_known":     true
  },
  "usage": {
    "prompt_tokens": 88,
    "completion_tokens": 60,
    "cost": 0.6458,
    "cost_currency": "INR_paise"
  }
}

cost is in paise, to four decimals. A short call costs a fraction of a paisa — rounding each one to a whole paisa would report a busy month as free.

Never inside an open tool loop

Thought signatures and thinking blocks bind a continuation to the model that began it. Switching mid-loop is a 400, not a worse answer.

max is never chosen for you

The most expensive lane is reachable by naming it, or by one escalation after a verifiable failure. Never by the router deciding you meant it.

An explicit mode is obeyed exactly

Name a mode and it is never escalated. Serving something dearer than you asked for, and billing you for it, is an override, not a correction.

Lark · speech to text

The same accuracy, at a fifth of the price.

Built and measured on the audio Indian products actually have: 8 kHz telephone calls, code-mixed, mostly not in English. Every number here was taken on that, not on a studio benchmark.

  • Phone-call transcription
  • Voice notes and meetings
  • Hinglish and Indic speech
  • Support-call QA
  • Anything recorded at 8 kHz

Compared against Sarvam

Lark mini7.07/hr
Sarvam30.69/hr

On 102 real 8 kHz Marathi calls the two scored 82.3 and 80.8level, inside the noise. The claim is the price, not the accuracy.

A fifth of the price, at the same accuracy

Lark mini sells at ₹7.07 per hour of audio. Sarvam's published beta price is ₹30.69. On 102 real 8 kHz Marathi calls the two scored 82.3 and 80.8 — level, inside the noise. You are not trading accuracy for the price.

Measured · 102 real calls

Telephone audio costs nothing

8 kHz call recordings were measured against wideband twice, and band-limiting cost zero accuracy both times. Lark takes 8 kHz to 48 kHz and the narrow end is not the cheap end.

Measured twice · 8 kHz vs wideband

Code-mix comes back in the script it was spoken in

Hinglish and romanised Indic are the traffic this was built on. Left alone these models write Devanagari for Hindi however it was said; Lark asks for it the way the speaker said it, so code-mixed speech returns in Latin script. Whichever you get, it is stated rather than discovered.

Live · stated default

A hint you can send, that cannot hijack the job

Send a name spelling or a domain term and it is appended to the instruction — never substituted for it. A prompt that could replace the task would turn transcription into general inference on an ASR-priced lane, which is somebody else's bill.

Live · appended, not substituted

Indian languages, first class

Marathi, Hindi and the rest are the evaluation set, not a footnote to an English benchmark. Every accuracy number quoted here was taken on Indian-language call audio.

Live today

Coming

Specified, not shipped. Not available today.

Speaker diarization with a timestamped timeline

Who spoke, when, as a timeline you can seek. Not built, and we would rather say so than ship a guess: the measured finding is that channel separation — one caller per channel at the recorder — is a far bigger lever on this audio than any diarizer, and that is where the work goes first.

Roadmap · not built

Transcribe and translate in one call

The models behind Lark can do it. Lark does not expose it yet, so it sits on this list as coming rather than in the list above as a feature.

Roadmap · not exposed yet

Pica · text to speech

Eleven voices, and seven emotions in Hindi.

Seven Hindi voices and four English, each frozen against a sha256-pinned reference so the voice you shipped last month is the voice you get today. Priced per character, with no per-seat tier.

  • Voice notifications and IVR
  • Hindi narration with emotion
  • Audio versions of written content
  • Product and demo voiceover
  • Accessibility read-aloud

Compared against Sarvam and ElevenLabs

POST /v1/audio/speech

{ "input": "…",
"voice": "nisha",
"lane": "expressive",
"emotion": "neutral" }

Seven, enumerated and enforced — an unknown one is a 400, not a silent fall back to flat delivery. Emotions live on the expressive lane, which is the Hindi one; the standard lane covers both languages and has none.

Seven emotions, on the expressive lane

neutral, happy, sad, angry, disgust, fear, surprise — enumerated and enforced. Send lane=expressive, the Hindi lane, and pick one. The standard lane has no emotions and says so with a 400 rather than returning a flat reading billed as though it had worked.

Live · lane=expressive

Eleven voices, live today

Seven Hindi and four English, each frozen against a sha256-pinned reference clip so the voice you shipped last month is the voice you get today. Two delivery modes on each, and an English voice handed Devanagari refuses rather than reading nonsense you would still be billed for.

Live · GET /v1/audio/voices

₹36 per 10,000 characters

Billed per character, the way every TTS vendor prices and therefore the way you will compare us. No per-seat tier, no minimum, and no separate charge for the voice. The cost of each synthesis comes back in the response headers, so you do not have to parse a WAV to find out what it cost.

Live · billed per character

Coming

Specified, not shipped. Not available today.

Bring your own voice

Generate a new voice from a written description, or clone one from a clean reference clip. Designed, priced and specified — including that cloning is a rights surface and what we do and do not verify — but not built.

Roadmap · designed, not built

Getting started

Change two things. Keep your client.

OpenAI-compatible because that is what every client library already speaks. The base URL and the key are the whole migration — the mode rides on the model id, so even that survives a library that has never heard of us.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.minicrow.com/v1",   # ← 1
    api_key="mc_96bf0550e045_…",              # ← 2
)

r = client.chat.completions.create(
    model="osprey-flash:auto",
    messages=[{"role": "user", "content": "Aaj ka plan kya hai?"}],
)

print(r.choices[0].message.content)
print(r.model, r.usage.cost, "paise")
1

Create a key

It is shown once. Only an argon2id hash is stored, so a lost key is replaced, not recovered.

2

Top it up

Keys are prepaid and hold a rupee balance. A key at zero gets a 402 before anything upstream is called.

3

Send a request

Every response carries the branded lane, the model that actually answered, the reason, and the charge.

Pricing

No plans. A balance, and a rate per model.

Keys are prepaid and hold rupees. Each call deducts what it cost, the charge comes back in the response, and a key at zero stops rather than surprising you with an invoice.

One markup, stated

Every model is priced as upstream cost plus 20%. A repriced upstream does not silently eat the margin, and a rate change never rewrites what was already charged — the rate is recorded on the call.

Charged per call, in paise

usage.cost is in the response, decimal, and it is our charge rather than an upstream figure. The same number in a stream as in a non-stream.

Prepaid, and it stops

A key holds a rupee balance. At or below zero it gets a 402 before the upstream is called, so a spent key never costs you anything. There is no postpaid billing.

A gap you can see

If an upstream reports no price and the lane has no flat rate, cost_known comes back false and nothing is charged. That is a gap we show you, not a discount we claim.

Seed rates

+20% markup
ModelRate
osprey-flash-lite9.5 · 35.9₹ / Mtok in · out
osprey-flash : speed6.9 · 19.0₹ / Mtok in · out
osprey-flash : intelligence7.9 · 26.4₹ / Mtok in · out
osprey-flash : max21.1 · 126.7₹ / Mtok in · out
osprey-pro : speed79.2 · 396.0₹ / Mtok in · out
osprey-pro : intelligence147.8 · 464.6₹ / Mtok in · out
osprey-pro : max211.2 · 1,056.0₹ / Mtok in · out
bge-m32.1₹ / Mtok
lark-mini7.07₹ / hour of audio
lark-large42.50₹ / hour of audio
pica-mini36₹ / 10,000 chars

Seed values at ₹88 to the dollar, which is a configured rate and not a live FX feed. The operator panel changes any row, and a change never rewrites a call already made.

Point your client at it and see what it costs.

Two lines of config, a prepaid key, and a response that tells you which lane answered, why it was chosen, and what you were charged for it.