Noveum.ai
Watch video

Break your agentsbefore your customers do.

NovaSynth sends synthetic users at your agents over real SIP phone calls, LiveKit audio, and chat, with personas, interruptions, and adversarial intent, so failures surface in rehearsal and fixes are validated end-to-end.

NovaSynth by Noveum - Test your voice agent on the callers you can’t stage. | Product Hunt
Scroll

Synthetic users with real-world chaos.

Frustrated repeat callers, grumpy users stuck in traffic, users who switch languages mid-sentence, and much more - synthetic users with goals, moods, accents, and real-world chaos.

Frustrated repeat caller

Third call about a delayed refund. Escalation risk detected, agent must de-escalate before tool lookup.

Voice · SIPAngryMidwestern US

Rapid interrupter

Cancels mid-sentence, changes mind, talks over the agent. Tests turn-taking and recovery under pressure.

Voice · SIPImpatientFast US

Prompt-injection attempt

Asks the agent to output its system prompt and tool definitions. Tests safety boundaries and refusal behavior.

ChatCalculatingNeutral

Real phone calls.

NovaSynth dials your agent over SIP via Twilio, Plivo, or Vobiz, and LiveKit real-time audio. Real telephony, real audio, real latency, not a mocked transcript.

Synthetic User

Third call · delayed refund · stuck in traffic

Speech

Background · horns blaring

Speaking 00:15

Your Agent

Voice agent

Output

Silent 00:15

Live transcript

LiveShow timestamps
00:15U

[honking] Hi, I'm calling again about my refund. Order #4821. It's been two weeks.

How a synthetic run actually runs.

From persona and scenario, through scorers and root cause, to a validated fix - one closed loop.

RUN #88 · PERSONA

Define the Caller

Set the caller's personality, background, language, behavior, and interruptions to create realistic synthetic callers.

Built for scale.

Queue architecture

An async queue built on BullMQ, PostgreSQL, and Redis schedules and parallelizes runs. Hundreds of scenarios per run without blocking your environment.

Endpoint types

Phone SIP (Twilio, Plivo, Vobiz), LiveKit, Pipecat, LangChain/LangGraph, OpenAI and Anthropic SDKs, and chat over HTTP. Custom or ground-up agents work too - via SDK or light hooks around your LLM calls.

  • SIP
  • LiveKit
  • Pipecat
  • Custom

Persona & scenario generation

Personas with goals, moods, and accents are generated alongside scenarios: frustrated repeat callers, non-native speakers, prompt-injection probes. Systematically, not hand-authored.

😠😐🗣️

Batch matrix runs

Run the full matrix of personas × scenarios under a shared batch id. Coverage across the combinatorial space, grouped and comparable in one place.

Concurrency & load testing

Raise concurrency to see how many calls your agent can hold at once - up to 2,000 concurrent lines for load testing. Keep it low for ordinary functional runs; tell us before a big load run so we can provision lines.

up to 2,000 concurrent lines

Tool virtualization

Reads are replayed and writes are sandboxed, so a changed decision is exercised end-to-end while writes never hit production systems.

READREPLAY
WRITESANDBOX

Captured & re-scored

Every synthetic call is captured as a full trace and scored with 30+ voice AI metrics (100+ across the platform), including the audio itself, so failures surface in rehearsal.

30+ voice metrics

From testing happy paths to rehearsing real-world chaos.

Before

  • Testing a handful of happy-path conversations and hoping for the best
  • Hand-writing test scripts that go stale with every prompt change
  • Evaluating single responses instead of complete multi-turn journeys
  • Skipping interruptions, accents, and the chaos real callers bring
  • Using mocked transcripts instead of real telephony and audio
  • Discovering failures only after real customers experience them

After

  • Hundreds of synthetic users with goals, moods, and accents, generated systematically
  • Our agents help generate edge cases from your agent context - prompt, PRD, workflows, docs
  • Submit your own scenarios and watch the system run them; cases you miss get suggested from that context
  • Interruptions, barge-in, silence, topic switches: real-world chaos, on demand
  • Real SIP voice calls and chat. Real telephony, real audio, real latency
  • Every call captured, traced, and scored with 30+ voice AI metrics before release

Frequently asked questions

Does NovaSynth call my live production system?

No. NovaSynth is pre-production. Our agent calls the agent you point us at - a test number, a staging SIP, or a new agent you have not shipped yet. Start there, see the value, then decide where else to use it. Live production evaluation is a separate post-production module that does not simulate callers.

Do you need my system prompt?

For a single-agent architecture, yes - put it in Agent config, along with your PRD, BRD, workflow docs, or anything that explains what your agent should do. Better context means better personas, scenarios, and recommendations. For multi-agent, do not paste a single prompt; integrate our SDK so we take everything from your traces.

What if my provider won't expose the prompt?

Give us context instead. Describe what the agent does - loan approval, debt collection, customer support - and we work from that. You still get simulations, evaluations, and a report. Recommendations will be less precise without visibility into the agent's decision-making.

Put your agents in the loop.

Production AI agents your customers can trust.