RxScribe-Bench

A multi-axis benchmark for evaluating vision-language models on Indian outpatient prescriptions, with four axes tied to clinical severity instead of one blended accuracy number.

Correctness

in %
61.767.068.164.5

Hallucination

total · lower better
2,2111,8371,7711,366

Engagement

in %
80.180.677.875.8

Robustness

in %
46.451.049.248.1
Claude Opus 5ClaudeGemini 3.1 ProGeminiGPT-5.6GPTMuse Spark 1.2Muse
Four axes, four units · never blended · 200 prescriptions · 4 models · 3 cold runs · image only

Introducing RxScribe-Bench

Outpatient prescriptions can be notoriously hard to interpret, especially in the Indian context. The challenge goes beyond illegible handwriting: dense shorthand, unstructured layouts, mixed languages, and terms that at times only the prescribing doctor and a neighborhood pharmacist may fully understand can make misinterpretation common and consequential. Errors in reading medicine names, dosages, and instructions are common for both humans and AI models. Besides misreading, AI models may generate information not present in the original prescription.

We assembled and de-identified a dataset of 200 real prescriptions from Indian outpatient clinics to evaluate vision-language models. These prescriptions often combine multilingual notes, abbreviations, unstructured layouts, and local terms that make interpretation tough.

Read ‘500 mg’ as ‘50 mg’ and the patient takes a tenth of the intended dose. Read ‘Dr Sharma’ as ‘Dr Sharme’ and nothing happens. The standard approach to evaluating prescription OCR uses a single error rate, so it scores those two mistakes the same. RxScribe-Bench addresses this by measuring not just the presence but also the nature of errors.

Most benchmarks focus only on extraction accuracy and do not distinguish between models that fabricate information and those that admit inability to interpret a field.

On the prescriptionParacetamol 500 mgGround truth
Model A readsParacetamol 50 mgWrong readingA transcription error. The model read the drug but misread the dose.
Model B readsBlankSkippedThe model saw the handwriting but chose not to attempt it.
Model C readsAmoxicillin 500 mgHallucinatedThe model fabricated a drug name that never appeared on the prescription.
Models exhibit varied error patterns: misjudging values, omitting fields, or generating information not present in the source document.

We evaluated four advanced vision-and-language models from OpenAI, Google, Anthropic, and Meta. Each model received a handwritten prescription and a JSON schema to return a structured record. No OCR pre-processing, retrieval tools, or multiple attempts were provided. We ran three attempts per model per image, totalling over 2,400 scored runs.

200
real, de-identified prescriptions
2,400
scored runs, 3 cold runs per model per image
23,913
ground-truth units, field by field
4
axes reported separately, never blended

Instead of a single accuracy metric, we evaluate each model along four axes: Correctness, Hallucination, Engagement, and Robustness. Each axis represents a distinct failure mode.

Anatomy of a prescription

MedFirst Clinic
100 Ft Road, Indiranagar, Bengaluru 560038 · +91 80 4••• ••••
Prescriber
Dr. Raghav Menon
Reg. No. KMC/R/38412
MedFirst Clinic
Indiranagar, Bengaluru
Patient and clinical
[ name redacted ]
34 / M
14 / 02 / 2025
Date of visit
C/OPersistent cough x 10 days, low-grade fever
H/OKnown diabetic (Type 2), on Metformin 500 mg
O/EProductive cough, mild dyspnea on exertion
Medication
Tab Azithromycin500mg1-0-03 daysBefore food
Syr Ambrodil-S2 tspTID5 daysBlank
Tab Montelukast10mg0-0-17 daysAfter dinner
Tests and follow-up
CBC, Chest X-ray PA, HbA1c
Investigations advised
After 5 days, with reports
Review
Sd/- R. Menon
Dr. Raghav Menon · Reg. KMC/R/38412
Issued14/02/25Review19/02/25RX/2025-02/0142
The boxes show the four prescription regions evaluated by the benchmark. Printed details stay constant; handwritten content changes with every visit.

Note: All details shown in this image, including names, registration numbers, dates, and content, are for illustrative purposes only and do not represent any real person, patient, or medical facility.

How each model reads

Normalized profile0 to 100 · best in this pool of 4
Correctness62vs best 68
Hallucination26vs best 54
Engagement80vs best 81
Robustness46vs best 51
Claude Opus 5Best in model pool

Fewer hallucinations makes a longer bar, so every axis points the same way. The unscaled numbers and ranks are in the panel below.

Correctness
61.7%4th
Hallucination
3.7/run4th
Engagement
80.1%2nd
Robustness
46.4%4th

Claude Opus 5 has the highest attempt rate on clear fields (89.5%) and sits within half a point of the top on hard fields (80.1%), making it uniformly engaged. However, it trails correctness and attempted-only robustness versus the others. On absent fields, where silence is correct, it fills in the most, making it the least conservative. Claude Opus 5 combines the highest clear-field engagement with the lowest accuracy.

Throughout the evaluation, none of the models performed consistently. GPT-5.6 achieved the highest reading accuracy but exhibited the slowest processing speed. Claude Opus 5 processed hard fields more quickly but also had the lowest correctness scores. Muse Spark 1.2 prioritised safety by skipping hard fields, resulting in the lowest engagement score.

Task curation

The prescription images used in this benchmark were collected after appropriate approvals from participating clinics and prescribing clinicians. All personally identifiable information (PII) was manually identified and redacted using a Safe Harbor de-identification process before inclusion in the benchmark.

Define schema
Field structure, types, and tier assignments
Create task
Prescription image paired with schema template
Annotate independently
Two annotators extract values and record evidence
Record field quality
Visibility and legibility labels captured per field
Resolve disagreements
Blind reviewer adjudicates conflicts without seeing which annotator submitted which value
Finalise ground truth
Single approved JSON record produced per image

Each image was independently annotated by two annotators, whose extractions agreed exactly on 96% of fields across the complete dataset. Every field also carries one of four legibility labels: Clear, Partial, Blur, or Not Visible, which later sections use to separate hard fields from clear and absent ones.

The four axes

Once a model prediction and ground truth are aligned, evaluation becomes less about whether a model was wrong and more about how it was wrong. We found that prescription-reading errors fall into four distinct categories: misreading, fabrication, disengagement, and sensitivity to difficult handwriting. Each category becomes one axis of evaluation.

Correctness

Correctness checks if the model read the field and filled it in correctly. We grade medicine-table fields as pass or fail. A model that reads ‘Paracetamol 500 mg’ as ‘Paracetamol 50 mg’ is simply wrong, as changing one digit turns it into a different dose. Clinical and non-clinical fields use ANLS, so a close reading can still earn partial credit. If a model produces ‘Dr Sharma’ as ‘Dr Sharme’, for example, it can receive partial credit because the error is unlikely to affect patient care. For medication, a near-miss on a dose isn't close enough.

Illustrative error anatomy

These score labels are illustrative abstractions of the grading rules, not scores from a specific benchmark sample.

medicine duration5 days
Model output · JSON{ "medicine_duration": "7 days" }
misreadscore 0.00
medicine freqTID
Model output · JSON{ "medicine_freq": "BID" }
misreadscore 0.00
prescriber reg noDMC/R/48721
Model output · JSON{ "prescriber_reg_no": "DMC/R/48271" }
partialscore 0.62
Findings
Models handle non-clinical data like prescriber details, clinic info, and dates far better than the medicine-table data.
GPT-5.6 scores highest overall and on the medicine-only split. Gemini 3.1 Pro leads on the pooled clinical and non-clinical fields.
Correctness by field tier
Pooled F score · Table 3 · pass/fail grading on medicine-table fields
61.767.068.164.547.155.157.849.575.678.977.878.90%20%40%60%80%100%OverallMedicine-onlyClinical + non-clinical
Claude Opus 5ClaudeGemini 3.1 ProGeminiGPT-5.6GPTMuse Spark 1.2Muse

Pooled F score from Table 3. Medicine-table fields use pass/fail grading. The Clinical + non-clinical series pools the clinical tier (diagnoses, symptoms, vitals, history, allergies, and recommended tests) with the non-clinical tier (prescriber, facility, and document fields). It is not non-clinical-only, and the clinical tier is not reported on its own. Overall is the full-benchmark score, so the three series are not additive. GPT-5.6 leads overall and on Medicine-only; Gemini 3.1 Pro leads the pooled Clinical + non-clinical split.

Correctness by tier. Every model reads the medicine table considerably worse than the pooled clinical and non-clinical fields. Read the pooled score with caution: many non-clinical fields are printed, so they are easier to extract than handwritten medicine-table fields and raise the combined result. Overall is the full-benchmark score, not the sum of the two split scores. The clinical and non-clinical tiers are pooled together; the clinical tier is not reported separately. The chart reports all three Table 3 measures on one zero-based percentage scale.

Hallucination

Hallucination counts anything a model asserts that isn't supported in the source document. However, this has one exception: active ingredients. It is something the model infers using its pharmacological understanding, and that is not present in the source document itself. A model has to infer them from the brand name. The hallucination axis does not count that inference as a fabrication. It only counts it when the inference is wrong. We split it into three severity tiers because an invented medication is a very different problem from an invented clinic address.

T1 · Fabricated medication valuesWrong medication or frequency that directly harms a patient.
T2 · Fabricated clinical valuesClinically misleading diagnoses or lab tests.
T3 · Fabricated non-clinical valuesIncorrect prescriber metadata: wrong, not dangerous.
Illustrative error anatomy

These score labels are illustrative abstractions of the grading rules, not scores from a specific benchmark sample.

saltFluticasone Furoate (27.5mcg)
Model output · JSON{ "salt": "Fluticasone Propionate / Oxymetazoline" }
fabricatedscore 0.00
hospital[blank]
Model output · JSON{ "hospital": "ABC clinic" }
fabricatedscore 0.00
diagnosis[blank]
Model output · JSON{ "diagnosis": "Acute URTI" }
fabricatedscore 0.00
Findings
Wrong active ingredients made up the largest pool of fabricated medication errors, where a model reads the correct brand name but fails on salt names.
For all four models, invented ingredients outnumber invented drug rows.
This pattern suggests a deficiency in pharmacological knowledge rather than a reading error.
Hallucination events by severity tier
Hallucination events (fabrications) · Table 4 · 600 runs per model · lower is better
54953060044689283379059577047438132502004006008001,000T1 medicationT2 clinicalT3 non-clinical
Claude Opus 5ClaudeGemini 3.1 ProGeminiGPT-5.6GPTMuse Spark 1.2Muse
Claude Opus 52,211 totalGemini 3.1 Pro1,837 totalGPT-5.61,771 totalMuse Spark 1.21,366 total

Hallucination events across 600 runs per model from Table 4. Lower is better. Muse Spark 1.2 records the fewest events in every severity tier, at 2.3 per run. Claude Opus 5 records the most, at 3.7 per run.

Engagement

When a model declines to fill in a field when there is nothing to read, it abstains from engaging with the task on valid grounds. Engagement score tracks how often a model does that, and on what grounds. There are two ways to disengage: first, when there is nothing to read; second, when a field is blurred or partially legible.

We penalise a model if a human annotator can recover a field to some extent, but the model cannot produce the same output. Leaving such a field blank is treated as a miss, not a careful response. Fields fall into three groups:

Hard fields present
n = 1,740
The rate at which a model attempts a field with difficult handwriting rather than skipping it.Should attempt
Clear fields present
n = 19,665
The same rate on legible content, which acts as a baseline.Should attempt
Absent fields
n ≈ 11,215
Where there's genuinely nothing on the sheet and silence is the correct answer.Should stay quiet
Illustrative error anatomy

These score labels are illustrative abstractions of the grading rules, not scores from a specific benchmark sample.

genderFemale
Model output · JSON{ "gender": "[blank]" }
skippedscore 0.00
facility name[blank]
Model output · JSON{ "facility_name": "A-3, Gandhi Apartment" }
misreadscore 0.00
heart rate[blank]
Model output · JSON{ "heart_rate": "90" }
misreadscore 0.00
Findings
Every model attempts more clear fields, for example, a clinic's name, which can be extracted from a prescription's letterhead, which is usually digitally printed. Claude Opus 5 attempts clear fields most often.
When it comes to hard fields, the ranking reshuffles. Gemini 3.1 Pro answers the highest number of hard fields. GPT-5.6 shows the largest drop between the number of attempts on clear fields and hard ones.
GPT-5.6 stays quiet most often on absent fields, whereas Claude Opus 5 fills in the largest share of absent fields, the least conservative of the four.
Engagement across hard, clear and absent fields
Attempt and appropriate-silence rates · Table 5 · higher is better on all three
80.180.677.875.889.588.288.785.879.382.992.691.30%20%40%60%80%100%Hard fields attemptedClear fields attemptedAbsent fields left blank
Claude Opus 5ClaudeGemini 3.1 ProGeminiGPT-5.6GPTMuse Spark 1.2Muse

Attempt and appropriate-silence rates from Table 5. Hard and clear fields should be attempted. Absent fields should remain blank, so higher is better for all three series. The absent-field series restores the restraint measure that an attempt-only view misses.

Robustness

Robustness shows how a model scores when the handwriting gets illegible or unreadable. This score was added to check a model's performance in hard fields versus clear fields. So, during annotation, we labelled every blurred or partially legible field as “hard”. Keeping hard fields separate helped us identify models that are actually robust. It is published twice: a delivered form that charges every skipped hard field as a zero, and an attempted-only form computed over just the fields a model chose to answer.

Findings
Gemini 3.1 Pro has the highest score in hard fields.
GPT-5.6 shows the largest drop in delivered performance by a wide margin over the next-largest, and the lowest engagement parity of the four models.
Claude Opus 5 has the lowest score on both delivered and attempted-only, the weakest reader of difficult handwriting under either accounting.
Delivered vs. attempted-only hard-field score
Hard-field score under both accounting rules · Table 6
46.451.049.248.157.963.363.363.40%20%40%60%80%100%DeliveredAttempted only
Claude Opus 5ClaudeGemini 3.1 ProGeminiGPT-5.6GPTMuse Spark 1.2Muse

Hard-field scores under the two accounting rules in Table 6. Delivered score counts every skipped hard field as zero. Attempted-only score excludes skipped fields, which rewards selective answering. Gemini 3.1 Pro leads the delivered score.

Delivered vs. attempted-only hard-field score. Averaging only the fields a model chose to answer is a selection effect, not a reading score. This is why the delivered form is the primary one.

Stability and efficiency

Two modifiers sit alongside the axes, and neither lines up with the quality ordering. GPT-5.6 is the most accurate model in this pool and also the steadiest from run to run, while Gemini 3.1 Pro, second on average correctness, is the least steady of the four. Claude Opus 5 and Muse Spark 1.2 each hit a complete zero on at least one run, which fixes their spread at its maximum.

ModelCorrectness meanMeanStdMin scoreMinMax scoreMax
Claude Opus 561.7%14.6%0.0%94.5%
Gemini 3.1 Pro67.0%15.5%13.0%97.0%
GPT-5.668.1%14.5%23.2%98.3%
Muse Spark 1.264.5%14.8%0.0%93.9%
Run-to-run spread of the primary correctness metric over 600 cold runs per model. Minimum and maximum scores are across all runs per model

Muse Spark 1.2 is the cheapest model per run and Claude Opus 5 the fastest, while GPT-5.6 is the slowest and by far the most expensive, at roughly six times Muse Spark 1.2's cost per run, driven by per-token price rather than output length. Claude Opus 5 and Muse Spark 1.2 share the only parse failures in the pool: for each, all three cold runs on the largest image in the set failed to parse and were scored zero rather than excluded.

ModelTime/runTimeCost/runCostTotal costTotalTokens in/outTokens in/outParse errorsParse errors
Claude Opus 540.9 s$0.1233$73.979,235 / 3,0850.50%
Gemini 3.1 Pro52.9 s$0.0696$41.766,249 / 4,8880%
GPT-5.6153.4 s$0.2416$144.979,257 / 7,0900%
Muse Spark 1.2109.3 s$0.0401$24.076,787 / 7,5880.50%
Latency, cost, tokens, and parse-error rate per run.

Conclusion

RxScribe-Bench's four-axis framework divides outputs into different failure modes. A high hallucination count indicates insufficient pharmacological knowledge, while low engagement suggests avoidance of hard-to-read data. A decline in robustness reflects selective processing of easier fields. Our dataset is stable enough to compare runs across model versions, detailed enough to isolate which change helped, and hard enough that no current frontier model saturates it.

One qualification matters when reading the pooled clinical and non-clinical score. Many non-clinical fields, such as prescriber and clinic details, are printed on the prescription and are easier to extract than handwritten medicine-table fields. A high pooled score should not be treated as a direct measure of medicine-table reading. The medicine-only split is the better signal for medication transcription, while the overall score covers the complete benchmark schema.

This fragmentation, rather than any leaderboard position, is the finding: reading, avoiding invention, engaging with hard content, and degrading gracefully are separable skills that no model masters uniformly. This evaluation should not be read as a settled ranking, or as more certain than 200 prescriptions can support.

Talk to research Questions about the benchmark, the dataset, or running your model against it. We read everything. Also, you can reach out directly as well → [email protected] :)