Skip to main content

Luna Voice Search Hardening Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.
Goal: Certify Luna as the only voice model and make voice instruction/knowledge retrieval speculative, multilingual, scoped, grounded, and reliable. Architecture: Luna emits an array of query variants. The backend batch-embeds and interleaves strict-scope instruction and knowledge searches, while a worker coordinator observes LiveKit’s preemptive LLM tool-call chunks and starts only the read-only search early. Search-first enforcement happens in the worker LLM node; workflow/output/harness fixes follow focused reproductions. Tech Stack: TypeScript, Vitest, Zod, LiveKit Agents 1.5.1, OpenAI embeddings, Kysely/Postgres/pgvector. Spec: docs/superpowers/specs/2026-08-26-luna-voice-search-hardening-design.md

Global Constraints

  • Voice evals and release gates target gpt-5.6-luna only; GPT-5.5 is deprecated.
  • Instruction and knowledge selections are independent strict allowlists; an empty selection means all content in that category is searchable.
  • Do not preload corpus content into the network-isolated worker.
  • Speculative execution is read-only knowledge search only.
  • Do not start, stop, or kill a shared backend or LiveKit worker.
  • Do not stage, commit, push, or create/update a PR without explicit git-write permission.

Task 1: Luna-only gate

Files:
  • Modify: .agents/skills/agent-evals/SKILL.md
  • Modify: backend/src/agent-evals/gate-runner-config.ts
  • Test: backend/src/agent-evals/gate-runner-config.spec.ts
  • Modify: backend/src/agent-evals/benches/foodics-voice-knowledge.bench.spec.ts
Interfaces:
  • Produces: default target list ['v2', 'voice-luna']; voice-luna has no deprecated comparison target.
  • Change the existing gate assertion to expect only v2 and voice-luna, and assert voice-gpt-5.5 is absent.
  • Run pnpm test src/agent-evals/gate-runner-config.spec.ts; expect the old target matrix assertion to fail.
  • Remove the GPT-5.5 target and comparison, make Luna the documented production voice target, and replace confusing voice-eval snapshot metadata defaults with Luna.
  • Rerun the focused spec; expect PASS.

Task 2: Multi-query search contract

Files:
  • Modify: backend/src/voice-runtime/dtos/knowledge-search.dto.ts
  • Modify: backend/src/voice-runtime/services/voice-knowledge-search.service.ts
  • Test: backend/src/voice-runtime/services/voice-knowledge-search.service.spec.ts
  • Modify: backend/voice-agent/src/api/schema.ts (generated)
  • Modify: backend/voice-agent/src/api/voice-backend-api.ts (generated-client consumer contract only if generation requires it)
  • Modify: backend/voice-agent/src/tools/knowledge.tools.ts
  • Test: backend/voice-agent/src/tools/knowledge.tools.spec.ts
Interfaces:
  • Consumes: queries: readonly [string, ...string[]], length 1-4, each query 1-500 characters.
  • Produces: one batched embedding call and one response containing strict-scope interleaved results.
  • Add DTO/tool tests proving query arrays are forwarded in order, deduplicated after normalization, and capped at four.
  • Run the two focused specs; expect schema/forwarding failures.
  • Change the wire DTO from query to queries; batch embeddings with input: queries.
  • Search both categories for every non-empty embedding, round-robin interleave per-query candidates by type:id:chunkIndex, then feed the combined candidates to the final ranker.
  • Generate the offline OpenAPI/client schema and rerun the focused specs; expect PASS.

Task 3: Category-aware relevance and pinning

Files:
  • Modify: backend/src/voice-runtime/services/voice-knowledge-ranking.utils.ts
  • Test: backend/src/voice-runtime/services/voice-knowledge-ranking.utils.spec.ts
  • Modify: backend/src/voice-runtime/dtos/knowledge-search.dto.ts
  • Modify: backend/src/voice-runtime/services/voice-knowledge-search.service.ts
  • Test: backend/src/voice-runtime/services/voice-knowledge-search.service.spec.ts
Interfaces:
  • Consumes: multiple query profiles and optional prior result IDs split by instruction/knowledge.
  • Produces: a maximum of five results, with both source categories represented when each has a relevant candidate; weak candidates produce no_match.
  • Add ranking tests for KDS positive-only queries, bilingual query union, per-category representation, authorized prior-ID pinning, and unrelated-corpus abstention.
  • Run the focused ranking/search specs; expect failures for missing multi-query scoring, pinning, and abstention.
  • Score each candidate against its best query profile, keep negated terms as penalties, pin authorized prior IDs without bypassing scope/segments, and add a conservative relevance cutoff based on vector distance plus lexical evidence.
  • Add telemetry for per-query candidate counts, best scores, selected IDs, and no-match reason.
  • Rerun focused specs; expect PASS.
Files:
  • Create: backend/voice-agent/src/knowledge-search-coordinator.ts
  • Create: backend/voice-agent/src/knowledge-search-coordinator.spec.ts
  • Modify: backend/voice-agent/src/tools/knowledge.tools.ts
  • Modify: backend/voice-agent/src/voice-agent.ts
  • Test: backend/voice-agent/src/tools/knowledge.tools.spec.ts
  • Test: backend/voice-agent/src/voice-agent.spec.ts
  • Modify: backend/voice-agent/src/index.ts
  • Modify: backend/src/agent-evals/voice-knowledge-eval.call.ts
  • Modify: backend/src/agent-evals/benches/moneygram-voice.bench.spec.ts
Interfaces:
  • Produces: VoiceKnowledgeSearchCoordinator.tools, observeLlmChunk(chunk), and exact-query promise reuse for search_knowledge only.
  • Add tests where an observed search_knowledge ChatChunk starts the backend request before tool.execute, and execution reuses the same promise exactly once.
  • Add tests proving malformed arguments, other tools, changed arguments, rejected searches, and cache eviction never trigger or reuse unsafe work.
  • Run coordinator/tool tests; expect missing-coordinator failures.
  • Implement a bounded per-call cache keyed by normalized validated args; no other tool name may enter it.
  • Change createAssistant to use LiveKit’s documented Agent.create({ llmNode }), delegate to Agent.default.llmNode, observe chunks without delaying/yield mutation, and preserve the original stream.
  • Wire the coordinator into production and in-process voice evals; rerun focused unit tests and expect PASS.

Task 5: Runtime search-first enforcement

Files:
  • Create: backend/voice-agent/src/knowledge-search-turn.utils.ts
  • Test: backend/voice-agent/src/knowledge-search-turn.utils.spec.ts
  • Modify: backend/voice-agent/src/voice-agent.ts
  • Test: backend/voice-agent/src/voice-agent.spec.ts
Interfaces:
  • Produces: VoiceKnowledgeSearchTurnUtils.requiresSearch(chatCtx): boolean and exact LiveKit function tool choice for search_knowledge only before the latest substantive caller turn has a search result.
  • Add literal chat-context tests for substantive policy questions, third-party questions, eight-to-twelve-digit references, greetings, yes/no acknowledgements, survey keypad answers, and post-search regeneration.
  • Run focused tests; expect failures because tool choice remains automatic.
  • Apply the exact function tool choice only on the first generation that requires search; preserve caller/model tool choice after a search result.
  • Rerun focused tests; expect PASS.

Task 6: Focused workflow state fixes

Files:
  • Modify only the reusable worker/backend tool-orchestration files identified by each failing focused test.
  • Add a colocated spec for every new function or state transition.
  • Preserve hard assertions in backend/src/agent-evals/benches/moneygram-voice.bench.spec.ts.
Interfaces:
  • Produces: typed workflow nextAction, requiredInput, and attemptsRemaining where supported; survey state owns question/answer/submission order.
  • Run Luna K=1 for failed-surname, caller-type-lock, card-troubleshooting, Spanish-survey, and survey-interruption scenarios without external workers; capture fresh artifacts.
  • For each reproducible failure, write one focused failing test at the tool/result boundary before changing production behavior.
  • Implement the smallest reusable state transition that fixes that failure; never key production logic by customer/org ID.
  • Rerun its focused unit test and scenario before moving to the next transition.

Task 7: Spoken output and eval harness

Files:
  • Modify: backend/src/agent-evals/voice-latency-eval.utils.ts
  • Test: backend/src/agent-evals/voice-latency-eval.utils.spec.ts
  • Modify the existing worker spoken-output utility selected by focused reproduction.
  • Add its colocated spec.
Interfaces:
  • Produces: terminal/tool-only latency classification and deterministic spoken formatting for telephone/reference numbers and uppercase initialisms.
  • Add failing tests for terminal end_call without a streamed sentence, phone/reference digit preservation, and initialism expansion.
  • Run focused specs; expect failures.
  • Classify terminal tool-only turns without manufacturing spoken latency, and normalize structured values before TTS.
  • Rerun focused specs; expect PASS.

Task 8: Review and Luna certification

Files:
  • Review every changed file with git diff --check and git diff.
  • Update: backend/src/agent-evals/research/foodics-voice-knowledge-search.md
  • Update: backend/src/agent-evals/research/moneygram-voice.md
  • Update: backend/src/agent-evals/research/ticketswap-voice-knowledge-search.md
Interfaces:
  • Produces: Luna-only GO/NO-GO report with 100% hard-scenario requirement.
  • Run touched unit/integration specs and pnpm tsgo; fix every failure.
  • Run Foodics, TicketSwap, and focused MoneyGram Luna K=1 in process; do not run worker cleanup commands.
  • Run the full Luna matrix at K=1; fix regressions.
  • Run focused failed scenarios at K=10.
  • Run pnpx bun src/agent-evals/gate-runner.ts --targets voice-luna --k 10 and report exact pass/fail counts, latency, cost, and artifact path.