Evaluation and methods for domain-grade models.

Open benchmarks, technical writing, and evaluation frameworks from the Zstate team. Agent trajectories, clinical reasoning, and regulated-domain preference data.

  • RxScribe-Bench: Structured Extraction from Indian Outpatient Prescriptions, with Hallucination and Abstention as First-Class Safety MetricsComing soon
  • The Missing Benchmark: Why AI's Next Crisis Will Be About Data, Not ModelsComing soon