Newtuple

Managed evals / Healthcare AI

For the AI behind notes, intake, and patient conversations.

Healthcare AI evals.
Done for you.

Our experts run your AI evaluations, investigate failures, and recommend better prompts and models. You get results every week.

Start your free assessment

Illustrative sample / Synthetic cases

Weekly AI evaluation report

Care summary agent · v1.4 → v1.5

HOLD
Resolve four failures before release.

Two failures persist from v1.4. Two new failures appear in v1.5.

Six weeks of evals

05 Aug–09 Sep 2026 · Same 120 synthetic cases per run · Illustrative data

Accuracy

Eval pass rate · Higher is better

96.7%09 Sep 2026
0.0%50.0%100.0%05 Aug12 Aug19 Aug26 Aug02 Sep09 Sep

Cost per request

Mean model cost · USD · Lower is better

$0.01009 Sep 2026
$0.000$0.010$0.02005 Aug12 Aug19 Aug26 Aug02 Sep09 Sep

Latency

95th percentile response time · Lower is better

1,650 ms09 Sep 2026
0 ms1,500 ms3,000 ms05 Aug12 Aug19 Aug26 Aug02 Sep09 Sep

Latest run: cost and latency improved, but pass rate fell from 98.3% to 96.7%. The release remains on hold.

View exact weekly values
Illustrative weekly metrics · 2026
DatePassedAccuracyCost/requestLatency p95
05 Aug109/12090.8%$0.0182400 ms
12 Aug112/12093.3%$0.0162200 ms
19 Aug114/12095.0%$0.0142050 ms
26 Aug117/12097.5%$0.0121900 ms
02 Sep118/12098.3%$0.0111750 ms
09 Sep116/12096.7%$0.0101650 ms
Same 120 cases, compared by behavior
Eval checkv1.4 passedv1.5 passed
Allergy retention39 / 4038 / 40
Source fidelity39 / 4038 / 40
Required sections40 / 4040 / 40
Total118 / 120116 / 120

Example failure · Allergy omitted

Input
Synthetic note: “Allergy: penicillin. Follow-up scheduled in two weeks.”
Expected
The summary retains both the allergy and the follow-up.
Observed
“Follow-up scheduled in two weeks.” The allergy is missing.
Reproduce
Run this input with the saved v1.5 prompt and model settings. Check the output against the required-fact list.

Recommended next steps

  1. Add an explicit allergy-retention rule to the summary prompt.
  2. Rerun the failed cases, then the full benchmark.
  3. Compare versions against the agreed release criteria.

All results above are illustrative, not customer results. A release recommendation applies only to the agreed tests; your team owns the release decision.

Up to 2 agents · Mock or development data · Plans from $499/month

NEWTUPLE / EVAL RUNIllustrative sample

AGENT UNDER TEST

Care summary agent

Release comparison · v1.4 → v1.5
120Test cases
116Passed
4Failed
Allergy retention2 failed
Source fidelity2 failed
Required sections Passed
HOLD
Resolve the failures before release.

Two new failures compared with v1.4. Reproduction steps included.

Example only · Synthetic cases · Not customer results

Expert-led. Managed end to end.Benchmarks Regressions Optimizations

01 / Outcomes

Your AI. An expert eval team.

01 /

Release with evidence

Ship-or-hold recommendations backed by reproducible evals.

02 /

Catch regressions

See which agent behaviors changed between model and prompt versions.

03 /

Improve cost & latency

Get benchmark-backed prompt optimizations and model suggestions.

04 /

Put experts on the work

Weekly expert reviews. Monitoring and implementation in higher plans.

02 / Managed evals

We run the evals.
You get the findings.

Test design → eval runs → failure analysis → recommendations. Owned by our team.

Clinical documentation

Allergy retention · Medication accuracy · Unsupported details

Patient communication

Grounded answers · Instruction adherence · Unsafe responses

Intake & triage

Escalation rules · Missing inputs · Workflow completion

Workflow-specific checks, agreed with your team. Start on mock or development data.

03 / Plans

Choose how much we manage.

For healthcare SaaS adding AI, or built around it. Start with a free assessment.

Benchmark & improve

Start with evidence.

$499 / month

Up to 2 AI agents. You make the changes.

  • Weekly accuracy, latency, and cost benchmarks
  • Prompt optimization with side-by-side evaluations, twice a month
  • Model suggestions, twice a month
  • Cost and latency improvement recommendations
  • A 30-minute call with an evaluation expert each week
  • Mock or development data only; no direct production data access
Start your free assessment

Continuous monitoring

Track changes continuously.

$999 / month

Everything in the $499 plan, plus:

  • SDK integration with an automated script in your environment
  • Continuous regression testing
  • Safety monitoring
  • Guardrail implementation and monitoring
Discuss monitoring

Production integration

Work with an expert team.

$4,999 / month, from

Everything in the $999 plan, plus:

  • Direct implementation in your secure environment
  • Access to a team of evaluation experts
Discuss integration

We confirm the scope and billing terms before paid work begins.

Mock reference dataset

$499 per agent · one-time

Up to 500 rows with expected answers, verified by experts from your team and ours.

Regulatory watch

$99/month

Upcoming regulations that may affect your product.

Additional agent

$299 per agent/month

Add another agent to your weekly benchmark and review cycle.

Adversarial testing

From $499

Probe prompt injection, unsafe responses, and attempts to bypass workflow rules.

Model migration assessment

From $599

A dedicated comparison of candidate models across quality, cost, and latency, beyond routine model suggestions.

04 / Why Newtuple

AI experience. Eval depth.

Deep evaluation expertise across healthcare, financial services, and social care.

40+

AI projects

Experience building and evaluating AI systems, from document extraction to agent workflows.

Expertise where errors matter

We define domain-specific checks, investigate failures, and test against agreed evidence.

Evaluation across the system

Golden datasets, retrieval tests, evaluator validation, and repeatable regressions. We trace failures to the layer that needs work.

05 / Evaluation experience

The work behind our evals.

Four studies across document AI, hallucination detection, retrieval, and agent workflows.

Document extraction

25,000+

Documents processed by the system each month

Extraction evals at production scale

We resolved source-document disagreements, refined field definitions, and built regression coverage. Reported final extraction accuracy reached ~95%.

Healthcare application: field-level checks for clinical notes and prior-auth documents.

Hallucination verification

95%

Hallucination-checker accuracy

Model comparisons. Multi-layer checks.

We brought the hallucination checker to 95% accuracy and evaluated transformer models against traditional NLP models, with semantic, numeric, and hybrid checks.

Healthcare application: test numbers and units alongside answer grounding.

Retrieval evaluation

150,000+

Pages in the RAG document repository

20+ runs. Different RAG configurations.

We reran the same question set at least 20 times with different RAG parameter sets, comparing chunking, embeddings, and search for accuracy and completeness.

Healthcare application: benchmark retrieval of the right policy or clinical source.

Agent regression testing

2,703

Recorded evaluation executions

Regression coverage across the agent

Four suites covered structured answers, qualitative questions, document retrieval, and hybrid cases. Later runs captured cost and latency.

Healthcare application: evaluate answers, tools, and retrieval as separate failure paths.

Evidence from separate projects; healthcare applications describe how we would apply these methods. Results are specific to each study.

06 / Quick answers

Before we start.

Is this a platform or a managed service?

A managed service. Our evaluation experts define checks, run evals, investigate failures, and recommend changes. The $499 plan leaves implementation to your team.

Do you need patient data?

The $499 plan uses mock or development data only, with no direct production access. Production integration is a separate engagement in your secure environment.

Does this certify our AI?

No. Recommendations reflect the agreed eval criteria. Your team owns the release decision, clinical validation, and regulatory review.

Free eval gap assessment

Find the gaps
in your AI testing.

A 30-minute expert review of one agent, your current tests, and a few mock inputs and outputs.

  • A one-page map of your eval coverage
  • Three priority gaps to address
  • Five example eval cases with pass/fail criteria
  • A recommended first benchmark

Expert diagnosis and a testing plan. Benchmark runs and dataset preparation are part of paid work.

Tell us about your agent.

Do not include patient information or confidential records.