Worker contract
LiveKit agent worker HTTP endpoints and client RPC contract for OptiMind sessions.
External agent workers (not in this repo) must talk to this API using X-Worker-Api-Key: $WORKER_API_KEY.
Dispatch metadata
Room/agent dispatch metadata is JSON matching:
{
organization_id: string
agent_id: string
agent_version_id: string
session_id: string
config: object
source: "web" | "campaign" | "reschedule" | "inbound" | "phone"
recording_enabled: boolean
channel: "WEB" | "SIP" | "PHONE"
direction: "NONE" | "INBOUND" | "OUTBOUND" | "WEB"
phone_number?: string | null
from_number?: string | null
sip_trunk_id?: string | null
livekit_sip_trunk_id?: string | null
campaign_id?: string | null
campaign_contact_id?: string | null
contact_metadata?: object
end_user?: { id: string, name: string } | null
}Every session belongs to an end user (end_user). Memories from that user's earlier sessions
are already merged into config.instructions under ## About the user.
Parse this from the LiveKit job metadata when the worker joins.
When config.avatar.enabled is true, config.avatar.provider_id ("anam" | "spatialreal")
selects the avatar plugin and config.avatar.external_avatar_id is the provider's avatar id.
For catalog avatars the API always fills provider_id from the catalog at dispatch time.
Required hooks
1. On session start (agent joins)
PATCH /api/internal/sessions/{session_id}/lifecycle
X-Worker-Api-Key: ...
{ "status": "ACTIVE", "livekitJobId": "...", "livekitWorkerId": "..." }If recording_enabled is true, start egress via the API (so an EgressJob row exists):
POST /api/internal/sessions/{session_id}/egress
X-Worker-Api-Key: ...
{ "audioOnly": true } # prefer true for SIP/PHONEDo not start LiveKit egress only from the worker without calling this endpoint.
2. Live model usage
Prefer session_usage_updated (session-level metrics_collected is deprecated):
POST /api/internal/sessions/{session_id}/events
X-Worker-Api-Key: ...
{
"eventType": "session_usage_updated",
"actor": "WORKER",
"payload": { /* AgentSessionUsage dump */ }
}Optionally also flush usage into the report endpoint periodically.
3. Tool calls
POST /api/internal/sessions/{session_id}/tool-calls
X-Worker-Api-Key: ...
{
"toolName": "end_call",
"arguments": {},
"result": "ok",
"status": "COMPLETED"
}4. On session end
Keep LiveKit text input enabled (lk.chat) so typed web chat enters session.history alongside spoken turns.
On AgentSession close, snapshot session.history.to_dict() / usage / metrics locally. Do not POST a stub report like { "metrics_flush": true } — that can overwrite a good livekitSessionReport and wipe transcript extraction.
Inside on_session_end:
report = ctx.make_session_report().to_dict()
# ensure chat_history.items present (fallback to close snapshot of session.history)
usage = report.get("usage") # list OR { model_usage: [...] }POST /api/internal/sessions/{session_id}/report
X-Worker-Api-Key: ...
{
"report": {
"chat_history": { "items": [ /* message / function_call / ... */ ] },
"events": [ /* AgentEvent dumps */ ],
"usage": [ /* or omit; prefer top-level usage */ ],
"...": "other SessionReport fields"
},
"usage": { "model_usage": [ /* LLM/STT/TTS summaries */ ] },
"metrics": [ /* optional plugin metric dumps */ ],
"collectedData": [
{ "key": "dob", "value": "1990-01-01", "label": "Date of birth", "fieldType": "string", "required": true }
],
"isFinal": true
}usage may be either an object ({ model_usage: [...] }) or a raw array of model usage rows.
collectedData is optional; when present, it is stored on AgentSession.metadata.collectedData (key → value).
Then:
PATCH /api/internal/sessions/{session_id}/lifecycle
X-Worker-Api-Key: ...
{ "status": "COMPLETED", "endReason": "COMPLETED" }Built-in web search
When config.tools_config.web_search is true, the worker attaches the model provider’s
native web-search tool (OpenAI WebSearch or Gemini GoogleSearch). No backend HTTP
endpoint is involved. Unsupported conversation providers ignore the flag.
End-user files
Push any file produced during the session (documents, images, summaries) for the session's end user:
POST /api/internal/sessions/{session_id}/files
X-Worker-Api-Key: ...
Content-Type: multipart/form-data
file=<binary> (required, max 25 MB)
name=<display name> (optional, defaults to the uploaded file name)Response: { file: { id, name, type, size, url } } (url is the storage key). Stored in S3 under
end-users/{org}/{endUser}/{session}/…, recorded as a SessionFile and appended to EndUser.files.
Returns 400 when the session has no end user.
Web participants can also upload (e.g. ID or face capture) with a LiveKit participant JWT:
POST /api/sessions/{session_id}/participant-files
Authorization: Bearer <livekit-participant-jwt>
Content-Type: multipart/form-data
file=<binary> (required, max 5 MB)
name=<display name> (optional)
participantToken=... (optional alternative to Authorization header)Token video.room must match the session's LiveKit room. Session must be QUEUED or ACTIVE.
Client ↔ agent LiveKit RPCs
add_context (candidate → agent)
When config.session_modalities.proctoring.enabled is true, the web client runs MediaPipe
checks and may call this RPC on the agent participant. OTP submit/resend also uses
add_context when session_modalities.otp_input.enabled is true.
Payload (JSON string):
{
state: string // spoken verbatim when action is "say"
action: "say" | "generate_reply" | "silent"
type: "multiple_people" | "additional_device" | "no_face" | "gaze_away"
| "id_captured" | "face_captured"
| "otp_submitted" | "otp_resend_requested"
details?: Record<string, unknown>
}proactive_response: falseforcesaction: "silent"for violation types.id_captured/face_capturedare sent withaction: "generate_reply"after a successful upload.otp_submittedincludesdetails.code; the worker/agent validates the code.otp_resend_requestedasks the agent to resend the OTP out-of-band.
start_id_capture (agent → candidate)
Opens the ID capture overlay. No payload required. Response: { "started": true }.
Registered only when session_modalities.proctoring.id_verification is true. See
ID verification.
start_face_capture (agent → candidate)
Opens the face capture overlay. No payload required. Response: { "started": true }.
Registered only when session_modalities.proctoring.face_verification is true. See
Face capture.
start_otp_input (agent → candidate)
Opens the OTP entry popover. Optional payload:
{ "length": 6, "hint": "Enter the code sent to your phone" }Response: { "started": true }.
Registered only when session_modalities.otp_input.enabled is true. See
OTP input.
Callbacks / reschedule
POST /api/internal/callbacks/schedule
X-Worker-Api-Key: ...
{
"sessionId": "...",
"scheduledAt": "2026-09-13T10:00:00+05:30",
"phoneNumber": "+15551234567",
"campaignId": "...",
"contactMetadata": {},
"source": "RESCHEDULE"
}Response: { ok, callbackId, contactId, nextAttemptAt, status }. Creates a CallbackSchedule row and, when a campaign contact is known, reschedules that contact for the dialer.
The report / terminal lifecycle handlers:
- merges into
livekitSessionReport(never replaces a full report with a metrics-only stub) - extracts
chat_history.items→Transcript+TranscriptSegment(spoken + typed chat) - upserts
SessionUsagefromusage/usage.model_usage - materializes
report.events→SessionEvent(agent.event.*) - materializes function call items →
ToolCallRecord - appends
metrics[]→SessionEvent(agent.metric.*) - stores
collectedData→AgentSession.metadata.collectedData - on
COMPLETED, updates the end user's memories from the transcript (once per session) - syncs linked
CampaignSessionoutcome / duration / transcript / recording from the agent session
Org-authenticated APIs
| Method | Path | Purpose |
|---|---|---|
POST | /api/sessions | Create AgentSession + LiveKit room/dispatch |
GET | /api/sessions | List sessions |
GET | /api/sessions/{id} | Session + transcript + usage + egress |
POST | /api/sessions/{id}/egress | Start room-composite egress → S3 |
POST | /api/sessions/{id}/end | Cancel session + delete LiveKit room |
POST | /api/sessions/{id}/participant-files | Participant JWT upload → SessionFile |
Webhooks
Configure LiveKit Cloud/project webhook to POST /api/webhooks/livekit.
Handled events:
egress_updated/egress_ended→ updateEgressJobstatus +fileUrlroom_finished→ complete openAgentSessionroom_started→ setlivekitRoomSid
Env
| Variable | Purpose |
|---|---|
WORKER_API_KEY | Worker auth for /internal/* |
AGENT_NAME | LiveKit agent name for dispatch |
S3_ACCESS_KEY_ID / S3_SECRET_ACCESS_KEY / S3_BUCKET_RECORDINGS / S3_REGION | Egress upload |
S3_ENDPOINT_URL (or S3_ENDPOINT) | Optional path-style endpoint (Supabase/MinIO/R2) |
S3_RECORDINGS_ACCESS_KEY_ID / S3_RECORDINGS_SECRET_ACCESS_KEY / S3_RECORDINGS_REGION | Optional overrides when knowledge + recordings use different keys |
LIVEKIT_URL / LIVEKIT_API_KEY / LIVEKIT_API_SECRET | LiveKit |
Recording object key (backend recordingFilepath ↔ worker recording_filepath):
{organizationId}/{sessionId}/{roomName}.mp4
→ s3://{S3_BUCKET_RECORDINGS}/...Reference worker
See optimind_operations/worker/main.py for the prior implementation. Port those hooks to call the endpoints above instead of writing egress without an EgressJob row, and rely on /report to materialize transcripts.