Noveum.ai
EV

Evaluation

Measure what your agents actually do.

NovaEval scores your real traffic with 112 calibrated evaluators across 15 categories: including the market’s only dedicated voice suite: and builds custom scorers from your product requirements.

112

calibrated scorers

15

scorer categories

36

dedicated voice scorers

app.noveum.ai: eval results
Noveum eval job score results

01: Capabilities

An instrument you can trust, not just a number.

112 scorers, 15 categories

Hallucination, RAG quality, answer correctness, tool use, coherence, safety: assembled from years of studying how language models fail systematically.

The only voice suite

36 dedicated voice scorers: mispronunciation, audio breakage, word error rate, talk-ratio, barge-in, dead-air, and full STT / LLM / TTS pipeline latency.

Panel of LLM judges

Disagreement-aware judging instead of a single model’s opinion: fewer hallucinated scores, calibrated verdicts.

Custom scorers from your PRD

Hand over your product requirements; instruction-driven scorers with templated placeholders: input_text, output_text, expected_output, retrieved_context, tool_calls: turn them into evaluation criteria, with scorers suggested from the traffic it sees.

Datasets from real traffic

ETL jobs convert production traces into living eval datasets, refreshed continuously: your logs become your benchmark.

Calibrated across versions

A 7/10 means the same thing after the model upgrade. Score distributions are monitored so your instrument never drifts silently.

02: The library

Eighteen categories. One calibration bar.

Every scorer reports a score and its reasoning, so you can audit the judge: not just trust it.

RAG EvaluationHallucinationAnswer QualityAgent EvaluationVoice LatencyAudio & TTS QualitySafetyBias DetectionConversationalAccuracy & MatchingContext AnalysisFormat ValidationNLP MetricsPanel JudgingMulti-ContextRelevanceG-EvalCustom
Hallucination detection with reasoning
Hallucination scorer with reasoning: live product

03: Under the hood

What every score is made of.

112 calibrated scorers, a dedicated voice suite, template-driven custom criteria, and datasets that build themselves: with reasoning attached to every verdict.

Scorer library

112 calibrated scorers across 15 categories: hallucination, RAG quality, answer correctness, tool use, coherence, safety, bias, format validation, and more: assembled from years of studying how language models fail.

Voice & audio suite

36 dedicated voice scorers: mispronunciation, audio breakage, word error rate (WER), talk-ratio, barge-in, dead-air, and STT / LLM / TTS pipeline latency: signals a transcript never shows.

Panel of judges

A panel of LLM judges with disagreement awareness replaces a single model’s opinion: fewer hallucinated scores, calibrated verdicts you can defend.

Custom scorers

Instruction-driven scorers with templated placeholders: input_text, output_text, expected_output, retrieved_context, tool_calls, and more: turn your product requirements into evaluation criteria.

Scorer output

Every scorer returns a score from 0–10, a passed boolean, free-text reasoning, and structured metadata: so you audit the judge, not just trust the number.

Datasets from traffic

Evaluation datasets are auto-built from production traces via ETL: your real logs become a living benchmark, refreshed continuously.

The loop

Part of one closed loop.

Evaluation is the control signal for everything else: NovaGuard enforces these scores in real time, and NovaPilot optimizes against them.

Next step

Put your agents in the loop

Production AI agents your customers can trust.

Trusted in production by enterprise and high-growth teams