# Webel (https://webel.ai/) Unleash your business. Webel is your company brain. Every customer and employee conversation compounds, and you book the value. Get started See how it works Your whole team One thread for the whole team. Your ops lead describes the problem, you set the guardrails, Webel does the work, in the same conversation, with every decision remembered. Work is a team sport; Webel is the AI built for it. See how teams work with Webel Build and ship real software. Most AI for business answers questions or writes drafts. Webel builds the actual software: it plans the work, writes the code, tests it, and ships it. You describe it; working software comes out the other side. Webel remembers your business. You never repeat yourself. Your suppliers, your preferences, your decisions, every build. It's time for your business knowledge to start compounding. Get the best model on every step. Webel picks the best model for each step: frontier when it matters, open source when it doesn't. The sharpest result at the lowest cost, and never locked to one vendor. This website was built by Webel. Our site, tools, and operations run on the same product we sell you. Start building. Describe what you need. Watch it ship. Get started Read the FAQ --- # Platform (https://webel.ai/platform) The platform How Webel works. You describe what you need in plain English. Webel plans the work, builds it, verifies it, ships it, and remembers what it learned. Here is the system underneath. The models The right model for every step. Frontier models for the hard problems, fast open-source models for the rest. Webel routes each step to the best fit and puts the cost of every call on your receipt. You never think about vendors. You get the sharpest result at the lowest price. Learn more → Memory It remembers your business. Every conversation, decision, and build sticks, and each fact traces back to where it came from. Six months in, Webel knows your suppliers, your stack, and your standards. You stop re-explaining. The work starts compounding. Learn more → Your team Your whole team, one thread. Ops describes the problem, engineering sets the guardrails, Webel does the work. Everyone sees what was asked, what was built, and who approved it. Nothing lives in a sidebar or a status meeting. Learn more → ◘ It checks its own work. Every change is tested before it counts as done. When a model gets something wrong, Webel catches it and fixes it before you ever see it. Learn more → 🔒 Your data stays yours. Your workspace is isolated. You control who on your team sees what. What you build and what you tell Webel stay yours. ▦ Built for builders, too. Webel is also a platform. Build your own software on the same engine that runs Webel. Build on Webel → Get started. Describe something. Watch it ship. Get started See pricing --- # Solutions (https://webel.ai/solutions) Solutions What you can build. Three teams, three businesses, three things they built by describing them. Wherever you're headed, Webel can build it. The proof Real deliveries, on the record. Every tool Webel builds ships as a verified pull request: cost, models, checks, and the timeline all on the record. This is a real workspace's delivery list: inventory alerts, a connection-pool fix, search ranking, webhook circuit breakers, billing reconciliation. The kinds of tools every business needs and most never get. 🚀 A SaaS startup A 6-person team, no engineers yet. They asked Webel for a customer onboarding portal that checks documents and routes approvals. Shipped their MVP without hiring a developer. 🏭 A 200-person manufacturer Their ops team tracks stock across spreadsheets and email. They asked Webel for an inventory tracker that reorders from suppliers, routes approvals, and flags low stock. Replaced the spreadsheets and stopped guessing when to reorder. 🏢 An enterprise team A procurement team at a 2,000-person company. They asked Webel for a vendor-approval workflow that used to take a week of email. Approvals now clear in a day. The pattern From a description to shipped software. Here's what they have in common, and what you'd have too. Customer portal Let customers check orders, update details, and message you, all in one place. Inventory and ordering Watch stock levels, get reorder alerts, and see what's running low before it sells out. Approvals and workflows Route requests to the right person, track decisions, and stop chasing people over email. Operations dashboard Pull your key numbers into one screen you can actually read each morning. Booking and scheduling Take bookings, avoid double-booking, and send reminders automatically. Anything you can describe If you can say it, Webel can usually build it. That's the point. Build your first tool. Describe it and watch it build. Get started How it works --- # Pricing (https://webel.ai/pricing) Pricing Compute at cost. Plus $0.25 per million tokens. Two lines: what the models cost, passed through untouched, and a $0.25 per-million-tokens fee on usage. No seats, no tiers, no markup. Pay for what you use, nothing else. Compute At cost You pay exactly what the model providers charge. We pass it through untouched and aggregate spend across customers for the best rate. Claude Sonnet 5: $2 in, $10 out per million tokens GLM 5.2: $1.40 in, $4.40 out DeepSeek V4 Flash: $0.44 in, $1.32 out Every other model, all at cost Platform fee $0.25/Mtok $0.25 per million tokens, flat. Same rate for input and output. That's the whole fee for the platform. No per-seat pricing No tiers, no feature gates One receipt for every dollar What that costs Three real examples. Token counts are illustrative. Your own receipt shows your real numbers, itemized line by line. A SaaS startup, one product A 6-person team, no engineers yet. About 4M output and 8M input tokens a month, routed across a couple of models. compute · Claude Sonnet 5 (60%) $34 compute · GLM 5.2 (40%) $11 platform fee · $0.25/Mtok $3.00 total $48 A 200-person company, several tools A manufacturer's ops team running inventory, approvals, and a dashboard. About 15M output and 30M input tokens a month, routed across a model mix. compute · Claude Sonnet 5 (50%) $105 compute · GLM 5.2 (30%) $32 compute · Claude Opus 4.8 (20%) $105 platform fee · $0.25/Mtok $11.25 total $253 A developer, on the API One developer shipping an AI product on the API. About 15M output and 75M input tokens a month, mostly cheap models, with a frontier model for the hard parts. compute · DeepSeek V4 Flash (80%) $42 compute · GLM 5.2 (15%) $26 compute · Claude Sonnet 5 (5%) $15 platform fee · $0.25/Mtok $22.50 total $106 Webel routes each step to the right model, so your compute spend reflects a mix, not a single frontier rate. Tokens are how AI usage is measured. What you don't pay. No per-seat fees. No tiers. No feature gates. No markup on compute. Two lines: compute at cost, and a flat $0.25 per million tokens. Questions Questions. How much does Webel cost? Compute passes through at cost, exactly what the providers charge. On top of that, a $0.25 per-million-tokens fee on usage. No per-seat pricing, no tiers, no markup. Do I pay per user? No. Your whole team uses Webel together, and you pay for what you use, not for how many people use it. Is there a markup on compute? No. Compute is passed through at cost. The only thing you pay Webel is the flat $0.25 per million tokens. What is the platform fee for? It's the fee for the platform itself: the memory, the model routing, the verification, the orchestration. One flat $0.25 per million tokens, same for input and output. Can I see what I'm spending? Yes. Every dollar is itemized and you get a receipt. Nothing is hidden. Start building. Compute at cost, plus $0.25 per million tokens. Get started Read the FAQ Buying for a team and want a walkthrough? Talk to sales . --- # Company (https://webel.ai/company) Company Webel builds Webel. Webel is a Seattle-based AI company founded in 2026 by Chris Spanton and Brian Spanton. We build software from a plain-English description, and our own site, tools, and operations run on the same product we sell you. Why we started Software should build itself. Building software was too hard for the businesses that need it most. Small teams were stuck with spreadsheets, or waiting on developers, or paying for tools that didn't fit. We thought it should be easier. So we built the thing that makes it easier. Every claim on this page is something Webel does for us first, every day. That's the proof, not a promise. Founders Webel loves to build. Chris Spanton Co-founder & CEO Builds the product and tells the story. Brian Spanton Co-founder & CTO Builds the engine underneath it. Previously Meta and Spring Health. How we build. At cost. Proven, not promised. Built by agents, directed by people. A small team points a large agentic workforce at a goal, and the leverage is the product. Build with us. Get started Talk to us --- # Developers (https://webel.ai/developers) Developers Build on Webel. One OpenAI-compatible endpoint, your own API key, and the same engine Webel runs on: model routing, durable memory, and per-key spend. Quickstart Five lines to your first completion. Create a key, then point any HTTP client at the endpoint. Full docs at webel.ai/docs . 1. Create a key Self-serve in the app: name it, set an optional dollar cap, copy it once. No sales call, no waiting. 2. Make one call Any HTTP client. OpenAI-compatible wire format, so your existing SDK works. Just change the base URL. 3. Read the receipt Tokens, exact cost, and the conversation id to continue the thread, on every response. curl https://api.webel.ai/v1/chat/completions \ -H "Authorization: Bearer $WEBEL_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "auto", "messages": [{ "role": "user", "content": "What is model routing?" }] }' Create an API key Read the quickstart The routing, visible Watch the router work. Every call is itemized: which model ran, how many tokens, exactly what it cost. model: "auto" is not a black box; it's a line item. This is the by-model breakdown from a live workspace. Model allocation Pin a model, or let Webel route. Static and dynamic, both through the same endpoint. You decide per request. Static: pin a model Set model to a specific model id and that exact model runs. { "model": "claude-sonnet-5", "messages": [{ "role": "user", "content": "..." }] } Dynamic: auto-select Set model to auto and Webel routes to the best fit. The response tells you which model ran. { "model": "auto", "messages": [{ "role": "user", "content": "..." }] } The response's usage.model_selected_by is "pinned" or "auto" , so you always know who chose the model. Response Everything you need, in one round trip. Tokens, cost, and the conversation id to continue it later. response · 200 { "id": "chatcmpl-…", "object": "chat.completion", "model": "claude-sonnet-5", "choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }], "usage": { "prompt_tokens": 120, "completion_tokens": 80, "total_tokens": 200, "cost_microusd": 1250000, "model_selected_by": "auto" }, "webel": { "conversation": "…", "turn": "…", "flavor": "api" } } Python Or from your language of choice. Plain HTTP, so it works anywhere. Continue a conversation by passing its conversation id back in. python · requests import requests resp = requests.post( "https://api.webel.ai/v1/chat/completions", headers={"Authorization": f"Bearer {key}"}, json={ "model": "auto", # or "claude-sonnet-5" to pin "messages": [{"role": "user", "content": "Summarize this."}], "stream": False, }, ) data = resp.json() print(data["choices"][0]["message"]["content"]) print(data["usage"]["model_selected_by"]) # "auto" or "pinned" print(data["usage"]["cost_microusd"]) # what this call cost print(data["webel"]["conversation"]) # id to continue this thread Streaming Server-sent events, standard shape. Set stream: true and read deltas as they arrive, ending with data: [DONE] . curl · stream: true curl https://api.webel.ai/v1/chat/completions \ -H "Authorization: Bearer $WEBEL_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "auto", "stream": true, "messages": [{ "role": "user", "content": "Write a haiku about shipping." }] }' Why build on it Things you don't get from a raw model API. 🔗 Durable conversations Not stateless. Pass a conversation id back and the thread keeps its full history, on the graph. 💸 Per-key spend Every response carries cost_microusd . Accumulate it per key. OpenAI and Anthropic don't give you this. 🧭 Model routing Pin a model or route dynamically. One call, every model, frontier and open source. Start building. The engine Webel runs on, behind one endpoint. Get your API key Explore the docs --- # API Docs (https://webel.ai/docs) Start here Overview Quickstart Guides Authentication Models Streaming Rate limits & spend caps Conversations & memory Reference API reference Errors Elsewhere Why Webel for devs llms.txt Developer docs The Webel API. One OpenAI-compatible endpoint backed by the engine Webel runs on: model routing across frontier and open-source models, durable conversations with memory, and per-key spend on every response. If you have used the OpenAI Chat Completions API, you already know the shape. Point your client at https://api.webel.ai/v1/chat/completions , present a wbl- key, and go. Create an API key Five-minute quickstart What you get OpenAI compatibility. The request and response follow the Chat Completions wire format, so existing OpenAI SDKs work: change the base URL to https://api.webel.ai/v1 and swap in your wbl- key. Durable conversations. Pass a conversation id back and the thread continues with its full history. State lives in your Webel room, not in your prompt window. Per-key spend. Every response carries cost_microusd plus remaining credit headroom. Neither OpenAI nor Anthropic gives you per-key cost. Model routing. Pin a specific model, or send "model": "auto" and let Webel route each request to the best fit. The response tells you which model ran and who chose it. List available ids with GET /v1/models . Base URL BASE https://api.webel.ai/v1 All endpoints below are relative to this base. The API is served over HTTPS only; plain-HTTP requests are refused. Where to go next New here? The quickstart takes you from key to first completion in five minutes. Wiring up auth? Read Authentication . Choosing a model? See Models : list live ids with GET /v1/models, pin or route. Want the full request/response contract? See the API reference . Tuning cost control? See Rate limits & spend caps . 💡 Building an AI answer engine or crawler? Machine-readable site content is at /llms.txt and /llms-full.txt ; the docs pages themselves are plain server-rendered HTML. Next: Quickstart → Create an API key --- # Quickstart (https://webel.ai/docs/quickstart) Start here Overview Quickstart Guides Authentication Models Streaming Rate limits & spend caps Conversations & memory Reference API reference Errors Elsewhere Why Webel for devs llms.txt Quickstart From key to first completion. Five minutes, no SDK required. You need to be an owner or admin of a Webel room with the public API enabled. 1. Create a key In Webel, open your room's API keys page at app.webel.ai/api-keys (also linked from Settings). Click Create key , give it a name, and optionally set a spend cap in dollars. The raw key ( wbl-… ) is shown once , at creation. Copy it then; only a hash is stored after that. If this is the first key in the room, you may be asked to accept the API Terms of Service. It is one click and never asked again. Keys expire after 90 days. Revoke and replace on your schedule; revocation is instant. A key acts as its own read-only service identity scoped to exactly one room. It cannot act as you, and it cannot touch another room. 2. Make your first call One POST. The response follows the OpenAI Chat Completions shape. curl · POST /v1/chat/completions curl https://api.webel.ai/v1/chat/completions \ -H "Authorization: Bearer $WEBEL_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "auto", "messages": [{ "role": "user", "content": "What is model routing?" }] }' The first call mints a durable conversation in your room. The reply arrives synchronously, typically within seconds. 3. Read the response response · 200 { "id": "chatcmpl-…", "object": "chat.completion", "model": "…", "choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }], "usage": { "prompt_tokens": 120, "completion_tokens": 80, "total_tokens": 200, "cost_microusd": 1250000, "model_selected_by": "auto" }, "webel": { "conversation": "…", "turn": "…", "flavor": "api" } } usage.cost_microusd is what the call cost in millionths of a dollar ( 1250000 = $1.25). usage.model_selected_by tells you whether the model was pinned by you or chosen by Webel's router. webel.conversation is the thread id: keep it to continue the conversation. 4. Continue the conversation Pass the conversation id back as conversation . The thread keeps its full history server-side, so each follow-up sends only the new user message. { "model": "auto", "conversation": "", "messages": [{ "role": "user", "content": "Now compare it with static routing." }] } Using an OpenAI SDK Because the wire format is OpenAI-compatible, existing SDKs work with two changes: the base URL and the key. python · openai sdk from openai import OpenAI client = OpenAI( api_key=os.environ["WEBEL_API_KEY"], # wbl-… base_url="https://api.webel.ai/v1", ) resp = client.chat.completions.create( model="auto", messages=[{"role": "user", "content": "Summarize this quarter."}], ) print(resp.choices[0].message.content) Webel-specific fields ride along in every response body under usage.model_selected_by and webel.* ; SDKs ignore unknown fields, so nothing breaks. Troubleshooting 401 unauthorized. Missing or malformed bearer header, or the key was revoked or expired (keys live 90 days). 403 the public API is disabled for this room. Ask the room owner to enable the public API for the room, then retry. 400 no user message in `messages`. The request must include at least one message with "role": "user" . 402 payment required. The key hit its spend cap. Raise it on the keys page or let the period reset. Full status-code and error-body contract: Errors . Next: Authentication → ← Overview --- # Authentication (https://webel.ai/docs/authentication) Start here Overview Quickstart Guides Authentication Models Streaming Rate limits & spend caps Conversations & memory Reference API reference Errors Elsewhere Why Webel for devs llms.txt Authentication API keys are Bearer tokens scoped to your room. Every request authenticates with an API key sent as a bearer token. A key is a self-contained service identity: it belongs to exactly one room, acts read-only within it, and never borrows a human's permissions. Create an API key The Authorization header Authorization: Bearer wbl-… Requests without the header (or with anything other than a valid wbl- key) get 401 Unauthorized . Keys are shown in full exactly once, at creation. The stored form is a SHA-256 hash; a lost key cannot be recovered. Revoke it and mint a new one. What a key is Room-scoped. A key is bound to the one room it was created in and can never act in another. If your account has several rooms, create one key per room. Read-only service identity. The key runs as its own principal inside the room: it reads what it needs to serve your requests (including conversation history) but cannot write tenant data. Not you. The key is not a member of the room and does not inherit any human's memberships or permissions. Spend bills to the room's normal funding sources. Individually limited. Rate limits and spend caps apply per key, so one integration cannot starve another. Lifetime and rotation Keys expire after 90 days . The keys page shows each key's status and expiry. Revocation is instant. The next request with a revoked key gets 401 Unauthorized . Rotate on your schedule: create the replacement, deploy it, then revoke the old key. Both can be live at once during the window. Every key's last-used time is tracked and visible on the keys page, so stale integrations are easy to spot before they break. Handling keys well Keep keys out of source control: load from an environment variable or secret manager (the examples here use $WEBEL_API_KEY ). Use one key per integration or environment (dev / staging / prod). Per-key spend caps then bound the blast radius of any single leak. Name keys for their purpose ("staging-bot", "support-agent") so an audit of the keys page reads like an inventory. A leaked key cannot be scoped down after the fact: revoke it immediately and mint a replacement. Terms Accepting the Webel API Terms of Service ( app.webel.ai/terms ) is a one-time step, recorded per room and never asked again. Acceptance happens in the app (the "Accept API Terms" action) and is not re-checked per key. API calls themselves require nothing beyond the bearer header. Note for browser testing: the wbl- bearer key only applies when there is no Webel session cookie on the request. If you're signed in to webel.ai in the same browser, the session cookie takes precedence and /v1 returns 401 even with a valid key. Test from a server or an incognito window. Next: Streaming → ← Quickstart --- # Models (https://webel.ai/docs/models) Start here Overview Quickstart Guides Authentication Models Streaming Rate limits & spend caps Conversations & memory Reference API reference Errors Elsewhere Why Webel for devs llms.txt Models Discover, list, and pin models. Webel routes across frontier and open-source models through one endpoint. You can let Webel choose ( model: "auto" ) or pin a specific model by id. There is also an API to list exactly which ids are available right now. Listing available models GET /v1/models returns the live roster of model ids your key can use, in the OpenAI list shape, so OpenAI SDKs' built-in models.list() works unchanged: curl https://api.webel.ai/v1/models \ -H "Authorization: Bearer $WEBEL_API_KEY" { "object": "list", "data": [ { "id": "claude-sonnet-5", "object": "model", "created": 1756000000, "owned_by": "anthropic", "label": "Claude Sonnet 5", "context_window": 200000, "supports_images": true, "supports_documents": true }, { "id": "zai/glm-5.2", "object": "model", "created": 1756000000, "owned_by": "zai", "supports_images": true, "supports_documents": false } ] } data[].id is what you pin. Pass it as model on chat completions . label , context_window , supports_images , supports_documents are Webel extensions (additive; standard clients ignore them). They mirror what the product's own model picker shows. owned_by is the provider operating the model (e.g. anthropic , openai , crusoe , deepseek ); webel when unspecified. The same authentication applies as every other endpoint: a valid bearer key, room public API enabled. Errors follow the standard error envelope . The list is sorted by id and is per-request stable; ids are lowercase with provider prefixes where routing requires them (e.g. zai/glm-5.2 ), or bare for first-party ids (e.g. claude-sonnet-5 ). Why this list is always accurate /v1/models is not a static page we update by hand. It reads the same live catalog the Webel product itself uses for its model picker, from the same graph. When a new model is added or an old one retired inside Webel, this endpoint reflects it on the next request. There is no second source of truth to drift out of date. That is also why these docs deliberately do not enumerate model ids: any static list here could go stale. The endpoint is the contract. Using an id: pin vs auto Pin: pass "model": "" : that exact model runs. The response echoes it back in model , and usage.model_selected_by is "pinned" . Auto: pass "model": "auto" (or omit) and Webel routes each request to the best fit. The response names the chosen model, and usage.model_selected_by is "auto" . See the developers overview for why you might prefer this. A pinned id that isn't currently selectable (retired, or never existed) fails fast with a bad-request error rather than silently substituting another model. Check /v1/models if that ever happens. Next: Streaming → ← Quickstart --- # Streaming (https://webel.ai/docs/streaming) Start here Overview Quickstart Guides Authentication Models Streaming Rate limits & spend caps Conversations & memory Reference API reference Errors Elsewhere Why Webel for devs llms.txt Streaming Server-sent events, OpenAI shape. Set "stream": true to receive tokens as they are generated. The event stream follows the OpenAI chunk format, ending with a usage-bearing final chunk and data: [DONE] . The request curl https://api.webel.ai/v1/chat/completions \ -H "Authorization: Bearer $WEBEL_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "auto", "stream": true, "messages": [{ "role": "user", "content": "Write a haiku about shipping." }] }' The response is Content-Type: text/event-stream . If the request itself is rejected (bad key, exhausted cap), you get a normal JSON error with the proper status code instead. The stream only opens after the turn is accepted. The frames Each frame is an SSE data: line carrying one JSON chunk. Content arrives as delta chunks: data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Tokens"},"finish_reason":null}]} data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" arrive"},"finish_reason":null}]} The final chunk carries no content. It closes the choice and delivers the full usage block plus the webel block (conversation id, spend headroom): data: {"id":"chatcmpl-…","object":"chat.completion.chunk","model":"…","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":18,"completion_tokens":24,"total_tokens":42,"cost_microusd":210000,"model_selected_by":"auto"},"webel":{"conversation":"…","turn":"…","flavor":"api"}} data: [DONE] Usage arrives in the last data chunk , before [DONE] , not in a separate trailing event. Cost and the conversation id ride the same final chunk via usage.cost_microusd and webel.conversation . Chunk object is chat.completion.chunk ; everything else matches the non-streaming contract. Errors mid-stream A failure after the stream has opened cannot change the status code, so it arrives as a final error frame followed by [DONE] : data: {"error":{"message":"timed out waiting for the reply","type":"stream_error"}} data: [DONE] Treat any frame containing an error object as terminal: stop consuming, surface the message, and retry the request if appropriate. Client notes OpenAI SDKs handle this automatically when stream=True ; unknown fields on chunks are ignored. Keep the connection alive. Events flush as they happen; there is no heartbeat padding frame. If your client disconnects mid-stream, the turn still completes server-side and its cost still accrues to the key. Next: Rate limits & spend caps → ← Authentication --- # Rate limits & spend caps (https://webel.ai/docs/rate-limits) Start here Overview Quickstart Guides Authentication Models Streaming Rate limits & spend caps Conversations & memory Reference API reference Errors Elsewhere Why Webel for devs llms.txt Rate limits & spend caps Per-key limits you can read off every response. Limits apply per key, never per account, never shared across integrations. Both the request budget and the dollar budget are visible in headers and response bodies, so your client can self-throttle instead of guessing. Request rate limits A key may sustain 2 requests/second with bursts up to 30 requests . Every response (success or failure) carries OpenAI/OpenRouter-style headers describing the key's bucket: Header Meaning x-ratelimit-limit-requests The burst capacity of the key's bucket (30 on current keys). x-ratelimit-remaining-requests Requests left in the burst right now. x-ratelimit-reset-requests Seconds until at least one request is available again. Present only when you are being throttled. Exceed the bucket and the request is refused with 429 Too Many Requests (the same three headers are still set). Read x-ratelimit-reset-requests , wait that long, retry once. The bucket refills continuously at 2 requests/second. Spend caps (per-key credit limit) Each key can carry a lifetime spend cap in dollars, set at creation and editable any time from the keys page. The cap bounds what a leaked or runaway integration can cost. Capped keys report headroom in every response: webel.limit_microusd (the cap) and webel.limit_remaining_microusd (what's left), both in millionths of a dollar. Uncapped keys omit these fields. When remaining headroom reaches zero, the next request fails with 402 Payment Required . A request that starts under the cap may finish slightly over it. Cost is known only after the model call, so the check happens before the request runs. Budget for one request of overshoot. Raising a cap takes effect immediately; no new key or redeploy needed. Seeing your spend You never need to poll for cost: usage.cost_microusd in every response is that call's exact cost: sum it client-side for live dashboards. The keys page shows accumulated charges per key alongside its cap and last-used time. Compute passes through at cost. Webel adds no markup to API usage (see pricing ). 💡 Hard-budget pattern: set the key's cap as your true ceiling, then alert client-side when webel.limit_remaining_microusd crosses a threshold (say 20%). The server-side cap catches what your alert misses. Next: Conversations & memory → ← Streaming --- # Conversations & memory (https://webel.ai/docs/conversations) Start here Overview Quickstart Guides Authentication Models Streaming Rate limits & spend caps Conversations & memory Reference API reference Errors Elsewhere Why Webel for devs llms.txt Conversations & memory Threads that live on the graph. Webel conversations are durable objects in your room, not stateless prompt exchanges. The API creates and continues them like any other Webel conversation. History, context, and memory carry over without you resending anything. Minting vs continuing Omit conversation : each request mints a fresh conversation. The id comes back in webel.conversation . Pass conversation : the request appends to that thread server-side, with its full history in context. Continuing is also cheaper: you send only the newest user message while the model still sees everything before it. There is no separate "fetch history" call. The room is the source of truth. { "model": "auto", "conversation": "1834708065435648", "messages": [{ "role": "user", "content": "Go on." }] } What the server remembers The full transcript of every turn in the thread (yours and the assistant's), in order, automatically included as context. Room memory. The engine's durable memory of your room (decisions, preferences, facts it has learned) informs replies the same way it does in the product. The room persona acts as the system prompt; your room's configured instructions apply to API turns too. Because threads are ordinary Webel conversations, your team can open one in the app, read exactly what the integration said, and even step in. API work is never a black box. Where API threads live API conversations are stamped flavor: "api" . They appear in your room's conversation picker under a dedicated filter, so integrations never pollute the human thread list by accident. The default view hides them; switch the filter to see API traffic alongside everything else. Notes for multi-turn design The messages array must contain at least one user message; its last user message becomes the new turn. Earlier entries in the array do not replace the stored history. The thread's own history always wins. A conversation id from another room is rejected. Keys act only inside their own room. Thread ids are stable; store them against your users or jobs to resume weeks later at full fidelity. Next: API reference → ← Rate limits & spend caps --- # API reference (https://webel.ai/docs/api-reference) Start here Overview Quickstart Guides Authentication Streaming Rate limits & spend caps Conversations & memory Reference API reference Errors Elsewhere Why Webel for devs llms.txt API reference POST /v1/chat/completions. Create a model response for a chat conversation. OpenAI Chat Completions-compatible; Webel extensions ride in the usage and webel blocks and never break existing clients. POST https://api.webel.ai/v1/chat/completions Authenticates with Authorization: Bearer wbl-… (see Authentication ). The reply is computed by running one full turn through the Webel engine (routing, moderation, and durable conversation state included), then returned synchronously (or streamed; see Streaming ). Request body Field Type Notes messages required array A list of messages comprising the conversation so far, each {"role", "content"} . Must include at least one user message: its last user entry becomes the new turn. System-role entries are accepted but not applied. System behavior comes from your room's persona. content is a string, or, for image and PDF input, an array of content parts (see Multimodal input ). model optional string ID of the model to use. Send a specific model id to pin it exactly, or omit / send "auto" to let Webel's router pick per request. If the API key has a pinned model, "auto" resolves to that pin; otherwise the router chooses. Default: "auto" . conversation optional string Webel conversation id to append this turn to. Omit to mint a fresh thread (id returned as webel.conversation ). See Conversations & memory . room optional string The room to run in. A key is scoped to exactly one room; if supplied, it must match that room. Usually safe to omit. stream optional boolean If true , partial deltas are sent as server-sent events. Default false . See Streaming . temperature optional number Accepted for wire compatibility. Not applied yet. max_tokens optional integer Accepted for wire compatibility. Not applied yet. ⚠️ v1 note: temperature , max_tokens , and client-supplied system prompts are accepted so existing SDK code runs unmodified, but they do not change generation yet. Pinning behavior-critical integrations should not rely on them. Multimodal input: images & documents A user message's content may be an array of typed parts instead of a plain string. Image and document parts are routed into the same engine path as product attachments, so the model must be vision- or document-capable (see GET /v1/models 's supports_images / supports_documents , which now describe a capability this endpoint actually accepts). { "messages": [ {"role": "user", "content": [ {"type": "text", "text": "What does this document say about refunds?"}, {"type": "image_url", "image_url": {"url": "data:image/png;base64,…"}}, {"type": "file", "file": {"filename": "policy.pdf", "file_data": "base64…"}} ]} ] } Images : {"type":"image_url","image_url":{"url":"data:image/;base64,…"}} . Inline data: URLs only (remote http(s) URLs are not fetched). Accepted types: image/png , image/jpeg , image/webp . Max 5 images, 6 MiB each. The OpenAI detail hint is accepted and ignored. Documents : {"type":"file","file":{"filename":"…","file_data":"…"}} . file_data is bare base64 or a data:;base64,… URL; the media type is taken from an explicit media_type , else the data URL, else the filename extension. Accepted types: application/pdf , text/plain , text/markdown . Max 10 documents, 10 MiB each. Reclassification : a document type sent through image_url (e.g. a PDF) is reclassified to a document; the media type is authoritative, never a rejection. Refusals : an unsupported part type, an over-allowlist media type, or an over-limit attachment refuses the whole request with a 400 naming the offending part/file. Response body { "id": "chatcmpl-1834708065435648", "object": "chat.completion", "created": 1756051200, "model": "claude-sonnet-5", "choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }], "usage": { "prompt_tokens": 120, "completion_tokens": 80, "total_tokens": 200, "cost_microusd": 1250000, "model_selected_by": "auto" }, "webel": { "conversation": "1834708065435648", "turn": "1834708065441792", "flavor": "api", "limit_microusd": 10000000, "limit_remaining_microusd": 8750000 } } Field Type Notes id string Completion id, derived from the turn. object string chat.completion (non-streaming) or chat.completion.chunk (streaming frames). created integer Unix timestamp of response creation. model string The model that actually ran. When you pinned, it echoes your pin; when auto-routed, it names the chosen model. choices[].message object The assistant reply ( role + content ). choices[].finish_reason string stop : natural completion or the reply wait timed out cleanly. usage.prompt_tokens usage.completion_tokens usage.total_tokens integers Token accounting for the call. usage.cost_microusd integer This call's cost in millionths of a dollar ( 1250000 = $1.25). Compute passes through at cost, with no markup. usage.model_selected_by string pinned when you chose the model, auto when Webel's router did. webel.conversation string The thread id. Pass it back as conversation to continue. webel.turn string The reply turn's id within the thread. webel.flavor string Always api for API-created threads. webel.limit_microusd webel.limit_remaining_microusd integers Present only on capped keys: the cap and remaining headroom, in millionths of a dollar. Every response also carries rate-limit headers. See Rate limits & spend caps . Model selection Pin: "model": "" runs exactly that model. Unknown or unavailable ids are refused with a bad-request error rather than silently substituted. Auto: "model": "auto" (or omitted) lets Webel route each request to the best fit across frontier and open-source models. If the API key has a pinned model, "auto" resolves to that pin instead; the pin is set on the key (not per request). The chosen model is returned in model , and usage.model_selected_by confirms who chose. Model ids follow the provider-prefixed convention used across the platform (e.g. claude-sonnet-5 ). Available ids are those selectable in your room. Moderation Prompts and completions pass through content classification. A flagged input is refused before any spend accrues; a flagged completion is blocked and logged rather than returned (fail-closed). Refusals surface as errors. See Errors . For multimodal input, the text parts, the extracted text of every document (PDFs included), and the content of every image are classified before posting, so a flagged input is refused before any spend accrues. Image screening covers the sexual, self-harm, and violence categories; child-safety ( sexual/minors ) in images is not covered by the classifier and relies on the separately-managed hash-match control (a roadmap item). Next: Errors → ← Conversations & memory --- # Errors (https://webel.ai/docs/errors) Start here Overview Quickstart Guides Authentication Models Streaming Rate limits & spend caps Conversations & memory Reference API reference Errors Elsewhere Why Webel for devs llms.txt Errors Standard shapes, actionable messages. Errors use the OpenAI error envelope, with HTTP status codes you already know. The message is written to be actionable: it names what to fix. Error shape { "error": { "message": "rate limit exceeded", "type": "invalid_request_error", "code": 429 } } message : a human-readable explanation of what went wrong and, where possible, what to do about it. type : the error family. Currently invalid_request_error for all request-level failures and stream_error for mid-stream failures (see Streaming ). code : the HTTP status code, mirrored in the body for convenience. Status codes Status Meaning Typical cause / fix 400 Bad Request Malformed JSON; no user message in messages ; unknown or unavailable model id; bad conversation or room value. Fix the request body. 401 Unauthorized Missing/malformed bearer header, or the key is revoked or expired. Check the header; rotate the key if needed. 402 Payment Required The key's spend cap is exhausted. Raise the cap on the keys page or switch keys. 403 Forbidden The public API is disabled for this room. A room owner/admin can enable it in room settings. 405 Method Not Allowed The endpoint only accepts POST. 413 Request Entity Too Large The request body exceeds the 200 MiB limit, typically a multimodal payload over the per-turn ceilings (5 images × 6 MiB, 10 documents × 10 MiB, base64-encoded, see Multimodal input ). Trim or split the payload and resend; the request is refused before any spend accrues. 429 Too Many Requests Per-key rate limit exceeded. Read x-ratelimit-reset-requests , wait that many seconds, retry once. 500 Internal Server Error Something failed on our side. Safe to retry with backoff. If it persists, contact support. Retrying safely Retry with backoff: 429 and 5xx . Honor x-ratelimit-reset-requests on 429 rather than fixed sleeps. Do not retry blind: 400 , 401 , 402 , 403 , 413 will fail again until something changes on your side. Timeouts: replies wait up to ~3 minutes (180 seconds) before the endpoint reports a timeout error. The wait ceiling is operator-tunable on our side (180s is the default). If your client times out first, the turn may still complete server-side and accrue cost. Prefer a client timeout at or above 180s, or use streaming so progress is visible. OpenAI SDK retry policies work unchanged: the codes and headers match what they expect. Back to Quickstart → ← API reference --- # FAQ (https://webel.ai/faq) FAQ Questions. Common questions about Webel. What is Webel? Webel is AI that plans, builds, and ships real software for your business. You and your team direct it in plain English. How much does Webel cost? Compute passes through at cost, exactly what the providers charge. On top of that, a $0.25 per-million-tokens fee on usage. No per-seat pricing, no tiers, no markup. You pay for what you use. See example costs → Who is Webel for? Founders, small and medium businesses, and startups who want software without hiring developers. If you were priced out of custom software, Webel is for you. Do I need to know how to code? No. You describe what you need in plain English and Webel does the building. You don't write a line. What can Webel build? Internal tools, customer portals, dashboards, trackers, booking systems, and more. If you can describe it, Webel can usually build it. How does Webel keep my data separate? Your workspace is isolated from everyone else, and you control who on your team sees what. What you build and what you tell Webel stay yours. Does Webel learn from my data? Webel remembers your business so it can keep working for you. It does not train on your content or share it. How is Webel different from hiring a developer or agency? You describe the software in plain English and it ships in hours, not months. No quotes, no ticket queues, no waiting. And it keeps working on it after it ships. Still curious? Ask us anything Get started --- # Contact (https://webel.ai/contact) Contact Talk to us. Questions, a demo, or just to say hi. We read everything. Talk to sales Get hands-on help. Planning a rollout, want a guided demo, or need a hand sizing Webel for your team? Tell us where you are and a real person replies. Self-serve works too: start building right now . Request a demo Get more information Partner with us Name Work email Company Your role optional Phone optional Team size optional Select 1-10 11-50 51-200 201-1000 1000+ What would you like to see? optional Request a demo We reply to every request ourselves. Name Email Company optional What can we help with? Send message We reply to every request ourselves. Name Work email Company Your role optional Company website optional Type of partnership Select Referral / reseller Technology integration Co-marketing Investment Other Tell us about the partnership you have in mind Send it over Our partnerships team reads every note. Other ways Other ways to reach us. ✉️ Email Write to hello@webel.ai and we'll get back to you. Something in the product not working? File a support ticket and get a reference number. 🚀 Start building Describe what you need and start building. Get started → 💼 Partnerships Referral, integration, co-marketing, or something else entirely. Tell us what you're making. Partner with us → --- # Support (https://webel.ai/support) Support Answers first. A ticket when you need one. Most answers live in the docs. When you need a human, file a ticket, keep the reference number, and we reply by email. 📚 Self-serve The docs cover the API, models, rate limits, and spend caps. The FAQ covers pricing and the product itself. Read the docs → Common questions → 🎫 File a ticket Describe the problem, pick a severity, and get a reference number right away. A human reads every ticket. File a ticket ↓ 🟢 Platform status Live component status and the incident record, checked straight from the same signals our own monitoring uses. Check platform status → Ticket form File a ticket. Fill this in and your email app opens with everything structured and ready to send. Your reference number appears the moment you submit. Name Email Room or workspace id optional Severity SEV1 Production is down or data is at risk. SEV2 A core feature is broken, no workaround. SEV3 Questions, requests, and small bugs. What happened Attachment optional A composed email cannot carry files. Send the ticket first, then reply to your own ticket email with the file attached and the reference number in the subject. Please fill in the required fields: name, a valid email, a severity, and a description. Open email with my ticket Prefer plain email? Write to support@webel.ai . Ticket WEBEL-XXXXX Your email app should have opened with the ticket below, already addressed to support@webel.ai. Send it and we are on it. If nothing opened, copy the text and email it yourself. Open email again Copy ticket text Keep the reference. When we reply it comes from support@webel.ai with the same reference in the subject. Reply to that thread to add context. How tickets work Nothing gets lost. The reference number is the tracking key for your whole issue. It rides in the subject line so every reply stays in one place. Severity Use it when First response SEV1 Production is down, or data is at risk. Same business day SEV2 A core feature is broken and there is no workaround. One business day SEV3 Everything else: questions, requests, small bugs. Two business days First response is a reply from a person at support@webel.ai. The times above are targets in business days, Pacific time. Every ticket gets a reply, including the small ones. New to Webel? Most answers are a page away, and building starts in minutes. Get started Read the FAQ --- # Platform status (https://webel.ai/status) Platform status Is Webel up right now? Live platform health, checked from your browser every 30 seconds. If something looks broken, this page shows what we know, and the ticket form is one click away. Checking… Talking to the live status feed. Components API The engine behind app.webel.ai and the Webel API. … Database The durable store behind every conversation. … App (app.webel.ai) The web app you log into. … Recent history No incidents reported. This log is curated by hand. When something breaks, we write down what happened, how long it lasted, and how it was resolved. This page reflects what the platform's own health checks see, live. It does not yet track historical uptime percentages. We'd rather publish a short true record than a long smooth chart. How this works Your browser checks the same health signals our own monitoring uses. Nothing cached, nothing staged. The page itself is static and hosted separately from the platform, so it stays up even when other parts do not. Serving build Something wrong? Check the FAQ first, then file a ticket. Pick SEV1 if production is down and it goes to the front of the queue. File a ticket Read the FAQ → --- # Model routing (https://webel.ai/model-routing) Platform The right model for the job. Every model has strengths and blind spots. Webel routes each step to the best fit and puts the price of every call on your receipt. How it works You don't pick. It does. Webel decides which model handles each step, watches the result, and adjusts. Hard reasoning gets frontier models. Routine steps run on fast, cheap ones. The routing is visible in your spend, not buried in a bill. ↓ Cheaper where it can be A step routed to GLM 5.2 ($1.40 in, $4.40 out per million tokens) instead of Claude Sonnet 5 ($2 in, $10 out) does the same job for less. Simple work shouldn't pay frontier prices. ↑ Sharper where it needs to be Hard problems get the strongest model available. The good stuff goes where it matters. ◆ Never locked in Webel works across frontier and open source. When a better model ships, your work just gets better. No migration, no renegotiation. The right model, automatically. Describe the job. Webel routes it. Get started How Webel works --- # Memory (https://webel.ai/memory) Platform Webel remembers your business. Not your prompts. Your business: the suppliers, the decisions, last month's incident, the reason behind the rule. It all sticks, and it all traces back to its source. How it works Facts, not chat history. Webel distills what it learns into durable, sourced facts. The DB pool cap is 40 after the August 2 incident. Acme's webhook endpoint is circuit-broken. Each fact links to the conversation it came from, so you can always ask why. ↗ It compounds Every build starts from everything before it. Month three is faster than month one, and it never starts over. 📁 No more re-explaining You told it once. Your suppliers, your preferences, and your rules are remembered exactly, not paraphrased. 🔒 Yours, not ours What Webel remembers lives in your workspace. You control it, and it goes when you go. Put your business knowledge to work. It's time for it to start compounding. Get started How Webel works --- # Teamwork (https://webel.ai/teamwork) Platform Your whole team, one thread. Founder, ops, engineering. Everyone points Webel at the same goal, in plain English, in one place. build -> verify -> ship --> The loop Describe. Watch it work. Ship. An ops lead describes the problem in plain English. Webel investigates, writes the fix, and shows its work: reading the config, editing the code, running the checks. Your engineer watches every step and weighs in when it matters. Nothing gets lost in a ticket queue, a spec, or a game of telephone. The ask goes straight from the person who knows to the thing that builds. Who owns what Everyone can direct it. You decide what ships. Nobody on your team needs to be technical to point Webel at a problem. And nobody loses control: follow-ups land as approval cards, work is attributed to the person who asked for it, and every change is verified before it counts as done. It adds up Every thread makes the next one smarter. What your team tells Webel sticks. The mitigation from last month's incident, who owns the pool, what the customer prefers: it's all remembered, traced to the conversation it came from, and applied to the next build. Your team compounds instead of repeating itself. 👥 Everyone can use it You don't need to be technical to direct Webel. The whole team talks to it in plain English. 🔀 Nothing lost in handoffs The ask goes straight from the person who knows to the thing that builds. No queue, no translation, no drift. 🧭 One source of truth Every decision, build, and change lives in one thread your whole team can see. Your whole team, together. Point everyone at the same goal. Get started How Webel works --- # Self-correction (https://webel.ai/self-correction) Platform It checks its own work. Models get things wrong. Webel tests what it builds, catches its own mistakes, and fixes them before they reach you. How it works Verified before it counts as done. Every change runs the checks before it ships. This delivery went from plan to merged with four of four checks passing, and the whole trail is on the record: who directed it, which models built it, what it cost, minute by minute. ✓ Checked before it ships Every change is tested and verified before it counts as done. Nothing ships blind. 🔍 A trail you can read What it found, what it fixed, and why: every delivery keeps the record. It's a receipt, not a black box. 🛡 Fewer surprises Mistakes get caught in the loop, not by your customers. And when a check goes red, Webel stops, diagnoses, and repairs before calling it done. Software that checks its own work. So you don't have to. Get started How Webel works