CHERNYX BENCHMARK™
Can we trust medical AI outside the lab?
Medical AI systems can demonstrate impressive performance in controlled evaluations. But clinical environments are different. Different patients. Different scanners. Different protocols. Different workflows.
Chernyx Benchmark explores how to independently evaluate what happens beyond the original benchmark.
“A benchmark can tell us how an AI performed.
But can it tell us whether that AI can be trusted in the real world?
Performance is only one part of the picture. A clinically useful evaluation also needs to consider robustness, external validation, safety, fairness, clinical utility, governance, and what happens when the environment changes.
The gap between benchmark and bedside
Performance can change drastically when patient populations, scanners, acquisition protocols, image quality, clinical workflows, or model versions change.
Development Dataset
Cleaned, curated, and annotated.
Published Results
High performance on the hold-out set.
Independent Evaluation
Vendor-neutral robustness checks.
Local Clinical Environment
Different populations, different scanners.
Real-World Monitoring
Continuous drift detection.
What Chernyx is exploring
CHERNYX
BENCHMARK
Clinical Performance
Standard accuracy, sensitivity, specificity on curated data.
Robustness
How does the system behave when real-world conditions are less than ideal?
Generalizability
Does performance hold across different environments and populations?
Safety
What happens when the system is wrong? Failure mode analysis.
Fairness
Evaluating bias across demographic and technical subpopulations.
Clinical Utility
Does it actually help the people using it in a real workflow?
Traceability
Tracking data provenance and model versioning through the lifecycle.
Governance & Compliance
Mapping technical evidence to regulatory expectations.
Research Methodology
Our framework explores a reproducible, vendor-neutral methodology for independently evaluating medical AI systems, moving from raw data to a comprehensive evidence profile.
Beyond a single score
We are exploring whether medical AI should be represented through a multidimensional evidence profile rather than a single headline accuracy number.
Chernyx Evidence Profile Prototype
A serious benchmark should test failure.
We are interested not only in where a system succeeds, but in understanding where and why it fails.
The failure modes matter as much as the headline score.
Performance is only part of trust.
The research will explore how technical evaluation can be connected with appropriate governance. Chernyx is researching how evaluation evidence can be organized and mapped against applicable governance and regulatory expectations.
Regulatory requirements vary by jurisdiction, intended use, and product classification.
Validation shouldn't stop
when the model goes live.
BEFORE DEPLOYMENT
Validate
DEPLOYMENT
Observe
REAL WORLD
Monitor
MODEL UPDATE
Re-evaluate
CONTINUOUS IMPROVEMENT
Learn
Where we're starting
Medical Imaging
Initial research focus: Radiology AI
First research question
"How can we create a reproducible, vendor-neutral methodology for independently evaluating a radiology AI system?"
Building on existing work
We are researching and learning from existing standards, frameworks, and initiatives. Chernyx is exploring how these pieces can be operationalized into an independent evidence framework.
We're looking for collaborators
Chernyx Benchmark is an R&D initiative. We are interested in collaborating with people who understand the problem from different sides.
Radiologists
Clinical perspective
Medical AI Researchers
Evaluation methodology
Biostatisticians
Statistical rigor
Healthcare & Regulatory Experts
Governance and implementation
Research Status
Research
Understanding the existing landscape
Methodology
Defining the benchmark
Prototype
Building the evaluation engine
Validation
Testing the methodology
Collaboration
Working with clinical and research partners
Medical AI doesn't just need better models.
It needs better evidence.
Chernyx Benchmark™ — A Chernyx AI Labs Research Initiative
Research phase. Methodology under development.
