True streaming
Server-sent events with keep-alive pings and hour-long connections, so long agent sessions never get cut off mid-task.
Spend them on Claude Fable 5 or any of 31 models — every plan reaches every model. One OpenAI-compatible endpoint, streaming, tool calling, and vision included.
# pip install openai
from openai import OpenAI
client = OpenAI(
api_key="cc_your_key",
base_url="https://freemodel.in/v1", ← only change
)
stream = client.chat.completions.create(
model="claude-opus-4.8",
messages=[{"role": "user", "content": "Hello!"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
Drop-in for the tools you already use
Any tool that accepts a custom OpenAI endpoint works — see setup guides
No SDK to learn, no migration. If your code talks to OpenAI, it already talks to us.
Sign up and generate an API key from your dashboard. Scope it to inference, models, or embeddings.
Set the base URL to our endpoint and paste the key. Everything else in your code stays the same.
Track every request, token, and dollar in the dashboard. Swap models without touching your integration.
Everything a production integration needs — not a thin proxy.
Server-sent events with keep-alive pings and hour-long connections, so long agent sessions never get cut off mid-task.
Full function-calling round trips, including parallel calls and streamed argument deltas — what coding agents run on.
Send screenshots, mockups, and diagrams as base64 or URLs to any vision-capable model, with per-model capability discovery.
Chain-of-thought output arrives in a separate field, so thinking never contaminates your answer. Token budgets handled for you.
Every request, token, latency figure, and cent — per key and per model. Real numbers, no estimates.
Top up directly with 300+ cryptocurrencies (USDT, BTC, ETH, SOL, TRX). Instant automated credit from anywhere in the world.
Lightweight and fast, or deeply capable — same API either way.
gemma-2-2b
Lightweight open model by Google. Fast and efficient for everyday tasks.
deepseek-v4-flash-0731
DeepSeek fast open model. Very affordable with solid reasoning, 304B params.
gpt-5.6-luna
OpenAI efficient model. Great value with strong reasoning, 290c/s speed.
deepseek-v4-pro-0813
DeepSeek open-source pro model. Excellent reasoning and coding, 1.6T params.
qwen3.8-27b
Alibaba Qwen compact open model. Efficient 27.8B params with solid capability.
seed-2.1-turbo
ByteDance Seed turbo model. Fast with solid reasoning and coding.
Spend your allowance on Claude Fable 5 or anything else in the catalog — all 31 models are on every plan. Bigger package, cheaper per token.
| Plan | Tokens / month | Price | Per 1M tokens |
|---|---|---|---|
| Free | 1M | Free | — |
| Starter | 30M | $1.20/mo | $0.04 |
| Basic | 100M | $3.50/mo | $0.035 |
| Pro Best value | 200M | $5.75/mo | $0.0288 |
| Scale | 500M | $11.50/mo | $0.023 |
| Unlimited No cap | Unlimited | $50/mo | — |
Run past your allowance and requests bill against your balance at each model's published rate — never a surprise invoice.
Start free with 1M tokens every month. Paid plans begin at $1.20/month for 30M tokens and scale up to unlimited tokens at $50/month. Every plan unlocks all 31 models including Claude Fable 5 — you are only buying token volume, never feature access.
No. Every plan, including the free tier, reaches all 31 models — Claude Fable 5 included. Plans differ only in monthly token volume and per-minute rate limits.
Yes — completely. Point any official OpenAI SDK (Python, Node, Go, .NET) at our base URL and pass your key. Request and response shapes match the OpenAI spec, including streaming chunks, tool calls, and error envelopes.
Anything that accepts a custom OpenAI-compatible endpoint: Kilo Code, Cline, Cursor, opencode, Claude Code, Continue, and Aider are all verified. Set the provider to "OpenAI Compatible", paste the base URL and key, and pick a model.
Tokens cover both your prompt and the model's reply, combined. Your plan allowance is spent first and resets monthly; anything beyond it bills against your account balance at each model's published per-token rate.
Yes. Streaming uses server-sent events with keep-alive pings and connections that stay open for up to an hour — suited to long agent sessions. Tool calling supports full round trips, parallel calls, and streamed argument deltas.
Requests fall back to your balance. With an empty balance you get HTTP 402 and a clear error before the request reaches a model, so you are never billed unexpectedly and never cut off mid-response. Top up or upgrade and you resume immediately.
We accept Cryptocurrencies including USDT (TRC-20) and TRX via direct on-chain verification. Payments are verified and credited automatically on-chain with zero middleman fees.
Create a key, change your base URL, and start shipping. Free tier included — no card required.