AGENT UNDER TEST
Care summary agent
Release comparison · v1.4 → v1.5Two new failures compared with v1.4. Reproduction steps included.
Example only · Synthetic cases · Not customer results
Managed evals / Healthcare AI
For the AI behind notes, intake, and patient conversations.
Our experts run your AI evaluations, investigate failures, and recommend better prompts and models. You get results every week.
Up to 2 agents · Mock or development data · Plans from $499/month
AGENT UNDER TEST
Two new failures compared with v1.4. Reproduction steps included.
Example only · Synthetic cases · Not customer results
01 / Outcomes
Ship-or-hold recommendations backed by reproducible evals.
See which agent behaviors changed between model and prompt versions.
Get benchmark-backed prompt optimizations and model suggestions.
Weekly expert reviews. Monitoring and implementation in higher plans.
02 / Managed evals
Test design → eval runs → failure analysis → recommendations. Owned by our team.
Allergy retention · Medication accuracy · Unsupported details
Grounded answers · Instruction adherence · Unsafe responses
Escalation rules · Missing inputs · Workflow completion
Workflow-specific checks, agreed with your team. Start on mock or development data.
03 / Plans
For healthcare SaaS adding AI, or built around it. Start with a free assessment.
Benchmark & improve
$499 / month
Up to 2 AI agents. You make the changes.
Continuous monitoring
$999 / month
Everything in the $499 plan, plus:
Production integration
$4,999 / month, from
Everything in the $999 plan, plus:
We confirm the scope and billing terms before paid work begins.
$499 per agent · one-time
Up to 500 rows with expected answers, verified by experts from your team and ours.
$99/month
Upcoming regulations that may affect your product.
$299 per agent/month
Add another agent to your weekly benchmark and review cycle.
From $499
Probe prompt injection, unsafe responses, and attempts to bypass workflow rules.
From $599
A dedicated comparison of candidate models across quality, cost, and latency, beyond routine model suggestions.
04 / Why Newtuple
Deep evaluation expertise across healthcare, financial services, and social care.
Experience building and evaluating AI systems, from document extraction to agent workflows.
We define domain-specific checks, investigate failures, and test against agreed evidence.
Golden datasets, retrieval tests, evaluator validation, and repeatable regressions. We trace failures to the layer that needs work.
05 / Evaluation experience
Four studies across document AI, hallucination detection, retrieval, and agent workflows.
Document extraction
Documents processed by the system each month
We resolved source-document disagreements, refined field definitions, and built regression coverage. Reported final extraction accuracy reached ~95%.
Healthcare application: field-level checks for clinical notes and prior-auth documents.
Hallucination verification
Hallucination-checker accuracy
We brought the hallucination checker to 95% accuracy and evaluated transformer models against traditional NLP models, with semantic, numeric, and hybrid checks.
Healthcare application: test numbers and units alongside answer grounding.
Retrieval evaluation
Pages in the RAG document repository
We reran the same question set at least 20 times with different RAG parameter sets, comparing chunking, embeddings, and search for accuracy and completeness.
Healthcare application: benchmark retrieval of the right policy or clinical source.
Agent regression testing
Recorded evaluation executions
Four suites covered structured answers, qualitative questions, document retrieval, and hybrid cases. Later runs captured cost and latency.
Healthcare application: evaluate answers, tools, and retrieval as separate failure paths.
Evidence from separate projects; healthcare applications describe how we would apply these methods. Results are specific to each study.
06 / Quick answers
A managed service. Our evaluation experts define checks, run evals, investigate failures, and recommend changes. The $499 plan leaves implementation to your team.
The $499 plan uses mock or development data only, with no direct production access. Production integration is a separate engagement in your secure environment.
No. Recommendations reflect the agreed eval criteria. Your team owns the release decision, clinical validation, and regulatory review.
Free eval gap assessment
A 30-minute expert review of one agent, your current tests, and a few mock inputs and outputs.
Expert diagnosis and a testing plan. Benchmark runs and dataset preparation are part of paid work.