Your capacity is live.Change the baseURL and carry on.
An OpenAI-compatible client. No new library, no code change beyond the base address.
Base address
https://fullcup.ai
Same route, same fields, same responses as the client you already use.
Your key
fc_live_••••••••••••••••••••••••
The key is shown once, in the provisioning email, and is stored nowhere. Lost the email? Reissue it — the previous key stops working that instant.
First call
curl https://fullcup.ai/v1/chat/completions \
-H "Authorization: Bearer fc_live_..." \
-H "Content-Type: application/json" \
-H "X-FullCup-Workload: onboarding" \
-d '{
"model": "@cf/openai/gpt-oss-120b",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 2048
}'
X-FullCup-Workload: the tag that lets a backer see what the capacity was spent on.
Token ceiling and reasoning models
Four of the five top models return empty content when max_tokens is low: reasoning eats the ceiling before the answer. If your first call comes back empty, raise max_tokens. It did not fail.
Models allowed in your class
Edge standard since 2025-09-11
- @cf/meta/llama-4-scout-17b-16e-instruct
- @cf/google/gemma-4-26b-a4b-it
- @cf/openai/gpt-oss-120b
- @cf/qwen/qwen3-30b-a3b-fp8
- @cf/zai-org/glm-5.3-flash
A model outside the class is refused before any capacity is spent. The list is dated: what counts is the version in force when the cup was funded.
Consumption is public by design.
Every request emits a signed receipt. Your public cup page shows tokens, tag and time, with no login, in about a minute. This is not fine print — it is why anyone trusts you with capacity.
See a consumption pagePOST /v1/chat/completions · GET /v1/models