OptimindOptimind By Sashflow
Documentation

Worker contract

LiveKit agent worker HTTP endpoints and client RPC contract for OptiMind sessions.

External agent workers (not in this repo) must talk to this API using X-Worker-Api-Key: $WORKER_API_KEY.

Dispatch metadata

Room/agent dispatch metadata is JSON matching:

{
  organization_id: string
  agent_id: string
  agent_version_id: string
  session_id: string
  config: object
  source: "web" | "campaign" | "reschedule" | "inbound" | "phone"
  recording_enabled: boolean
  channel: "WEB" | "SIP" | "PHONE"
  direction: "NONE" | "INBOUND" | "OUTBOUND" | "WEB"
  phone_number?: string | null
  from_number?: string | null
  sip_trunk_id?: string | null
  livekit_sip_trunk_id?: string | null
  campaign_id?: string | null
  campaign_contact_id?: string | null
  contact_metadata?: object
  end_user?: { id: string, name: string } | null
}

Every session belongs to an end user (end_user). Memories from that user's earlier sessions are already merged into config.instructions under ## About the user.

Parse this from the LiveKit job metadata when the worker joins.

When config.avatar.enabled is true, config.avatar.provider_id ("anam" | "spatialreal") selects the avatar plugin and config.avatar.external_avatar_id is the provider's avatar id. For catalog avatars the API always fills provider_id from the catalog at dispatch time.

Required hooks

1. On session start (agent joins)

PATCH /api/internal/sessions/{session_id}/lifecycle
X-Worker-Api-Key: ...
{ "status": "ACTIVE", "livekitJobId": "...", "livekitWorkerId": "..." }

If recording_enabled is true, start egress via the API (so an EgressJob row exists):

POST /api/internal/sessions/{session_id}/egress
X-Worker-Api-Key: ...
{ "audioOnly": true }   # prefer true for SIP/PHONE

Do not start LiveKit egress only from the worker without calling this endpoint.

2. Live model usage

Prefer session_usage_updated (session-level metrics_collected is deprecated):

POST /api/internal/sessions/{session_id}/events
X-Worker-Api-Key: ...
{
  "eventType": "session_usage_updated",
  "actor": "WORKER",
  "payload": { /* AgentSessionUsage dump */ }
}

Optionally also flush usage into the report endpoint periodically.

3. Tool calls

POST /api/internal/sessions/{session_id}/tool-calls
X-Worker-Api-Key: ...
{
  "toolName": "end_call",
  "arguments": {},
  "result": "ok",
  "status": "COMPLETED"
}

4. On session end

Keep LiveKit text input enabled (lk.chat) so typed web chat enters session.history alongside spoken turns.

On AgentSession close, snapshot session.history.to_dict() / usage / metrics locally. Do not POST a stub report like { "metrics_flush": true } — that can overwrite a good livekitSessionReport and wipe transcript extraction.

Inside on_session_end:

report = ctx.make_session_report().to_dict()
# ensure chat_history.items present (fallback to close snapshot of session.history)
usage = report.get("usage")  # list OR { model_usage: [...] }
POST /api/internal/sessions/{session_id}/report
X-Worker-Api-Key: ...
{
  "report": {
    "chat_history": { "items": [ /* message / function_call / ... */ ] },
    "events": [ /* AgentEvent dumps */ ],
    "usage": [ /* or omit; prefer top-level usage */ ],
    "...": "other SessionReport fields"
  },
  "usage": { "model_usage": [ /* LLM/STT/TTS summaries */ ] },
  "metrics": [ /* optional plugin metric dumps */ ],
  "collectedData": [
    { "key": "dob", "value": "1990-01-01", "label": "Date of birth", "fieldType": "string", "required": true }
  ],
  "isFinal": true
}

usage may be either an object ({ model_usage: [...] }) or a raw array of model usage rows.

collectedData is optional; when present, it is stored on AgentSession.metadata.collectedData (key → value).

Then:

PATCH /api/internal/sessions/{session_id}/lifecycle
X-Worker-Api-Key: ...
{ "status": "COMPLETED", "endReason": "COMPLETED" }

When config.tools_config.web_search is true, the worker attaches the model provider’s native web-search tool (OpenAI WebSearch or Gemini GoogleSearch). No backend HTTP endpoint is involved. Unsupported conversation providers ignore the flag.

End-user files

Push any file produced during the session (documents, images, summaries) for the session's end user:

POST /api/internal/sessions/{session_id}/files
X-Worker-Api-Key: ...
Content-Type: multipart/form-data

file=<binary>          (required, max 25 MB)
name=<display name>    (optional, defaults to the uploaded file name)

Response: { file: { id, name, type, size, url } } (url is the storage key). Stored in S3 under end-users/{org}/{endUser}/{session}/…, recorded as a SessionFile and appended to EndUser.files. Returns 400 when the session has no end user.

Web participants can also upload (e.g. ID or face capture) with a LiveKit participant JWT:

POST /api/sessions/{session_id}/participant-files
Authorization: Bearer <livekit-participant-jwt>
Content-Type: multipart/form-data

file=<binary>          (required, max 5 MB)
name=<display name>    (optional)
participantToken=...   (optional alternative to Authorization header)

Token video.room must match the session's LiveKit room. Session must be QUEUED or ACTIVE.

Client ↔ agent LiveKit RPCs

add_context (candidate → agent)

When config.session_modalities.proctoring.enabled is true, the web client runs MediaPipe checks and may call this RPC on the agent participant. OTP submit/resend also uses add_context when session_modalities.otp_input.enabled is true.

Payload (JSON string):

{
  state: string              // spoken verbatim when action is "say"
  action: "say" | "generate_reply" | "silent"
  type: "multiple_people" | "additional_device" | "no_face" | "gaze_away"
      | "id_captured" | "face_captured"
      | "otp_submitted" | "otp_resend_requested"
  details?: Record<string, unknown>
}
  • proactive_response: false forces action: "silent" for violation types.
  • id_captured / face_captured are sent with action: "generate_reply" after a successful upload.
  • otp_submitted includes details.code; the worker/agent validates the code.
  • otp_resend_requested asks the agent to resend the OTP out-of-band.

start_id_capture (agent → candidate)

Opens the ID capture overlay. No payload required. Response: { "started": true }.

Registered only when session_modalities.proctoring.id_verification is true. See ID verification.

start_face_capture (agent → candidate)

Opens the face capture overlay. No payload required. Response: { "started": true }.

Registered only when session_modalities.proctoring.face_verification is true. See Face capture.

start_otp_input (agent → candidate)

Opens the OTP entry popover. Optional payload:

{ "length": 6, "hint": "Enter the code sent to your phone" }

Response: { "started": true }.

Registered only when session_modalities.otp_input.enabled is true. See OTP input.

Callbacks / reschedule

POST /api/internal/callbacks/schedule
X-Worker-Api-Key: ...
{
  "sessionId": "...",
  "scheduledAt": "2026-09-13T10:00:00+05:30",
  "phoneNumber": "+15551234567",
  "campaignId": "...",
  "contactMetadata": {},
  "source": "RESCHEDULE"
}

Response: { ok, callbackId, contactId, nextAttemptAt, status }. Creates a CallbackSchedule row and, when a campaign contact is known, reschedules that contact for the dialer.

The report / terminal lifecycle handlers:

  • merges into livekitSessionReport (never replaces a full report with a metrics-only stub)
  • extracts chat_history.items → Transcript + TranscriptSegment (spoken + typed chat)
  • upserts SessionUsage from usage / usage.model_usage
  • materializes report.events → SessionEvent (agent.event.*)
  • materializes function call items → ToolCallRecord
  • appends metrics[] → SessionEvent (agent.metric.*)
  • stores collectedData → AgentSession.metadata.collectedData
  • on COMPLETED, updates the end user's memories from the transcript (once per session)
  • syncs linked CampaignSession outcome / duration / transcript / recording from the agent session

Org-authenticated APIs

MethodPathPurpose
POST/api/sessionsCreate AgentSession + LiveKit room/dispatch
GET/api/sessionsList sessions
GET/api/sessions/{id}Session + transcript + usage + egress
POST/api/sessions/{id}/egressStart room-composite egress → S3
POST/api/sessions/{id}/endCancel session + delete LiveKit room
POST/api/sessions/{id}/participant-filesParticipant JWT upload → SessionFile

Webhooks

Configure LiveKit Cloud/project webhook to POST /api/webhooks/livekit.

Handled events:

  • egress_updated / egress_ended → update EgressJob status + fileUrl
  • room_finished → complete open AgentSession
  • room_started → set livekitRoomSid

Env

VariablePurpose
WORKER_API_KEYWorker auth for /internal/*
AGENT_NAMELiveKit agent name for dispatch
S3_ACCESS_KEY_ID / S3_SECRET_ACCESS_KEY / S3_BUCKET_RECORDINGS / S3_REGIONEgress upload
S3_ENDPOINT_URL (or S3_ENDPOINT)Optional path-style endpoint (Supabase/MinIO/R2)
S3_RECORDINGS_ACCESS_KEY_ID / S3_RECORDINGS_SECRET_ACCESS_KEY / S3_RECORDINGS_REGIONOptional overrides when knowledge + recordings use different keys
LIVEKIT_URL / LIVEKIT_API_KEY / LIVEKIT_API_SECRETLiveKit

Recording object key (backend recordingFilepath ↔ worker recording_filepath):

{organizationId}/{sessionId}/{roomName}.mp4
→ s3://{S3_BUCKET_RECORDINGS}/...

Reference worker

See optimind_operations/worker/main.py for the prior implementation. Port those hooks to call the endpoints above instead of writing egress without an EgressJob row, and rely on /report to materialize transcripts.