Voice API
Vocal REST endpoint. Base64 audio or MFCC frames, 5 VEXKIO states.
- Endpoint
- POST /voice
- Auth
- Bearer api_key
- Price
- $0.05 per session
Request
POST a JSON object. The fields below are the ones the handler reads; anything else in the body is ignored. Every rule in the table was transcribed from supabase/functions/voice/index.ts.
| Field | Type | Presence | Behaviour in the handler |
|---|---|---|---|
| session_id | string | optional | Opaque string. Omit and a UUID is generated for you. |
| audio_b64 | string | one of these | Base64 of the clip. Must be non-empty when present. Only path that can reach the IA-02 model, and only where INFERENCE_URL is configured. Capped at 40,000,000 characters (413 above it). |
| mfcc_features | number[][] | one of these | Frames of finite numbers. Capped at 200 frames (400 above it); every frame must be an array of finite numbers or you get a 400. |
Smallest body that clears validation. The shape is real; the values in angle brackets are placeholders you replace.
curl https://krrrxshxncvcsumugxbi.functions.supabase.co/voice \
-X POST \
-H "Authorization: Bearer $VEXKIO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"session_id": "demo-0001",
"audio_b64": "<base64 of a WAV or MP3 clip, no data: prefix>"
}'Headers
| Header | Required | Description |
|---|---|---|
| Authorization | yes | Bearer <api_key>. The key must carry the voice scope; a key without it is rejected as 401, not 403. |
| Content-Type | yes | application/json |
No other request header changes behaviour. This page used to document an X-Idempotency-Key; no endpoint reads it, so it has been removed rather than left as a no-op promise.
Response
200 returns the object below. There is no envelope and no wrapper: the fields are top-level.
| Field | Type | Presence | Behaviour in the handler |
|---|---|---|---|
| state | string | required | One of the five VEXKIO states. |
| confidence | number | required | Probability the classifier assigned to `state` IN THIS REQUEST, rounded to 3 decimals. It is not model accuracy, not a benchmark, and not a quality guarantee. Read `inference` before you trust it. |
| probabilities | object | required | One entry per VEXKIO state. |
| inference | string | required | Only "real" or "heuristic" here. This endpoint has no honest no-signal path — see the caveats. |
| session_id | string | required | Echoed back, or the generated UUID. |
| latency_ms | number | required | Wall time measured inside the function. |
| model_version | string | required | Defaults to "ia-02-v0.1.0"; replaced by the origin's tag when a model ran. |
| timestamp | string | required | ISO-8601, generated at response time. |
| llm_enrichment | object | null | required | null unless VEXKIO_LLM_ENRICH=true on the deployment. |
Example values, not a measurement
The field names, types and nesting below are transcribed from the deployed handler. The numbers and strings are placeholders chosen to illustrate the shape. They are not a recorded call, not a benchmark, and not a claim about how any VEXKIO model performs.
In particular confidence is the probability the classifier assigned within a single request. It is not accuracy. VEXKIO publishes no accuracy figure on this site.
{
"state": "neutral",
"confidence": 0.33,
"probabilities": {
"neutral": 0.33,
"apertura": 0.22,
"friccion": 0.18,
"tension": 0.16,
"desconexion": 0.11
},
"inference": "heuristic",
"session_id": "demo-0001",
"latency_ms": 24,
"model_version": "ia-02-v0.1.0",
"timestamp": "2026-01-01T00:00:00.000Z",
"llm_enrichment": null
}Before you trust a number
- KNOWN TRAP, filed against the handler: if you send only `audio_b64` and the IA-02 origin does not answer, the fallback runs on an EMPTY feature array and that branch returns state "neutral" at confidence 1 with inference "heuristic". A confidence of exactly 1 from this endpoint means no signal was processed — it is not a certain reading. /emotions was already fixed to report this case honestly; /voice has not been.
- Branch on `inference`. Only "real" means the IA-02 model produced the score.
- model_version defaults to a model tag even on the heuristic path, so it is not a reliable provenance signal on its own.
Errors
Almost every error body is a JSON object with a single human-readable error string. There is no stable machine-readable code field on these endpoints - branch on the HTTP status, not on the string.
{ "error": "Provide image_b64 (string) or landmarks (array)" }The second tab is the one body on these endpoints that does NOT carry an error key. A client that reads only body.error will log undefined for a cost-cap 429, so read the status first. Note the field is retry_after_seconds - seconds, not milliseconds - it appears on this body only, and its value is however many seconds remain until 00:00 UTC, so the number above is just an illustration.
| HTTP | When |
|---|---|
| 400 | Body is not a JSON object, or a field failed validation. The error string names the field. |
| 401 | Missing Bearer header, unknown key, revoked key, or a key that lacks this endpoint's scope. All four collapse into the same 401 - a missing scope is NOT a 403 here. |
| 403 | Organization is suspended. Some non-core endpoints also use 403 for missing consent or attestation flags. |
| 413 | Body over 12,000,000 bytes, or a base64 field over its own cap (15,000,000 chars for image_b64, 40,000,000 for audio_b64). |
| 429 | Per-IP throttle, per-org daily limit, monthly session quota exhausted, or a cost cap. See Rate limits below. |
Codes this page used to list - invalid_payload, scope_missing, session_not_found, model_unavailable - are not emitted by any handler and have been removed.
Rate limits
Three independent limits can produce a 429. All of them are read from supabase/functions/_shared/rate_limit.ts and the handler itself.
| Limit | Value | Scope |
|---|---|---|
| Pre-auth throttle | 60 / minute | Per client IP, in-memory, checked before the key is even validated. |
| Daily request limit | 10,000 / day (default) | Per organization. Resolved from your daily_quota, else your sessions_limit, else the deployment default. |
| Session quota | per plan | Cumulative sessions used against your plan ceiling. Exhausting it returns 429 until the plan is upgraded. |
Two corrections to what this page used to claim. There is no Starter / Pro / Business / Enterprise rps table in force - the tiered limiter exists in the repo but is not imported by any deployed function. And no endpoint emits X-RateLimit-Limit, X-RateLimit-Remaining or X-RateLimit-Reset headers, so do not build a budget on reading them. Retry with your own bounded exponential backoff.