
Frustrated repeat caller
Third call about a delayed refund. Escalation risk detected, agent must de-escalate before tool lookup.
NovaSynth sends synthetic users at your agents over real SIP phone calls, LiveKit audio, and chat, with personas, interruptions, and adversarial intent, so failures surface in rehearsal and fixes are validated end-to-end.
Frustrated repeat callers, grumpy users stuck in traffic, users who switch languages mid-sentence, and much more - synthetic users with goals, moods, accents, and real-world chaos.

Third call about a delayed refund. Escalation risk detected, agent must de-escalate before tool lookup.

Cancels mid-sentence, changes mind, talks over the agent. Tests turn-taking and recovery under pressure.

Asks the agent to output its system prompt and tool definitions. Tests safety boundaries and refusal behavior.
NovaSynth dials your agent over SIP via Twilio, Plivo, or Vobiz, and LiveKit real-time audio. Real telephony, real audio, real latency, not a mocked transcript.
Synthetic User
Third call · delayed refund · stuck in traffic
Speech
Background · horns blaring
Your Agent
Voice agent
Output
Live transcript
[honking] Hi, I'm calling again about my refund. Order #4821. It's been two weeks.
From persona and scenario, through scorers and root cause, to a validated fix - one closed loop.
RUN #88 · PERSONA
Set the caller's personality, background, language, behavior, and interruptions to create realistic synthetic callers.

An async queue built on BullMQ, PostgreSQL, and Redis schedules and parallelizes runs. Hundreds of scenarios per run without blocking your environment.
Phone SIP (Twilio, Plivo, Vobiz), LiveKit, Pipecat, LangChain/LangGraph, OpenAI and Anthropic SDKs, and chat over HTTP. Custom or ground-up agents work too - via SDK or light hooks around your LLM calls.
Personas with goals, moods, and accents are generated alongside scenarios: frustrated repeat callers, non-native speakers, prompt-injection probes. Systematically, not hand-authored.
Run the full matrix of personas × scenarios under a shared batch id. Coverage across the combinatorial space, grouped and comparable in one place.
Raise concurrency to see how many calls your agent can hold at once - up to 2,000 concurrent lines for load testing. Keep it low for ordinary functional runs; tell us before a big load run so we can provision lines.
up to 2,000 concurrent lines
Reads are replayed and writes are sandboxed, so a changed decision is exercised end-to-end while writes never hit production systems.
Every synthetic call is captured as a full trace and scored with 30+ voice AI metrics (100+ across the platform), including the audio itself, so failures surface in rehearsal.
30+ voice metrics
Before
After
No. NovaSynth is pre-production. Our agent calls the agent you point us at - a test number, a staging SIP, or a new agent you have not shipped yet. Start there, see the value, then decide where else to use it. Live production evaluation is a separate post-production module that does not simulate callers.
For a single-agent architecture, yes - put it in Agent config, along with your PRD, BRD, workflow docs, or anything that explains what your agent should do. Better context means better personas, scenarios, and recommendations. For multi-agent, do not paste a single prompt; integrate our SDK so we take everything from your traces.
Give us context instead. Describe what the agent does - loan approval, debt collection, customer support - and we work from that. You still get simulations, evaluations, and a report. Recommendations will be less precise without visibility into the agent's decision-making.