Chernyx AI Labs Research

CHERNYX BENCHMARK™

Can we trust medical AI outside the lab?

Medical AI systems can demonstrate impressive performance in controlled evaluations. But clinical environments are different. Different patients. Different scanners. Different protocols. Different workflows.

Chernyx Benchmark explores how to independently evaluate what happens beyond the original benchmark.

The question we're exploring

A benchmark can tell us how an AI performed.
But can it tell us whether that AI can be trusted in the real world?

Performance is only one part of the picture. A clinically useful evaluation also needs to consider robustness, external validation, safety, fairness, clinical utility, governance, and what happens when the environment changes.

The gap between benchmark and bedside

Performance can change drastically when patient populations, scanners, acquisition protocols, image quality, clinical workflows, or model versions change.

Development Dataset

Cleaned, curated, and annotated.

Published Results

High performance on the hold-out set.

Independent Evaluation

Vendor-neutral robustness checks.

Local Clinical Environment

Different populations, different scanners.

Real-World Monitoring

Continuous drift detection.

Multidimensional Evaluation

What Chernyx is exploring

CHERNYX
BENCHMARK

Clinical Performance

Standard accuracy, sensitivity, specificity on curated data.

Robustness

How does the system behave when real-world conditions are less than ideal?

Generalizability

Does performance hold across different environments and populations?

Safety

What happens when the system is wrong? Failure mode analysis.

Fairness

Evaluating bias across demographic and technical subpopulations.

Clinical Utility

Does it actually help the people using it in a real workflow?

Traceability

Tracking data provenance and model versioning through the lifecycle.

Governance & Compliance

Mapping technical evidence to regulatory expectations.

Abstract Pipeline

Research Methodology

Our framework explores a reproducible, vendor-neutral methodology for independently evaluating medical AI systems, moving from raw data to a comprehensive evidence profile.

Medical Imaging Data
Reference Standard
AI System Under Evaluation
Independent Evaluation
Statistical Analysis
Robustness & Failure Analysis
Governance / Compliance Evidence
Chernyx Evidence Profile

Beyond a single score

We are exploring whether medical AI should be represented through a multidimensional evidence profile rather than a single headline accuracy number.

Illustrative concept — not a real model evaluation.

Chernyx Evidence Profile Prototype

Clinical Performance
Strong
Robustness
Under Evaluation
External Validation
Strong
Safety
Under Evaluation
Generalizability
Moderate Evidence
Governance
Evidence Required

A serious benchmark should test failure.

We are interested not only in where a system succeeds, but in understanding where and why it fails.

The failure modes matter as much as the headline score.

Image quality variation
Acquisition differences
Scanner variation
Out-of-distribution cases
Rare findings
Multiple simultaneous findings
Ambiguous cases
Population differences
Workflow changes
Model updates

Performance is only part of trust.

The research will explore how technical evaluation can be connected with appropriate governance. Chernyx is researching how evaluation evidence can be organized and mapped against applicable governance and regulatory expectations.

Regulatory requirements vary by jurisdiction, intended use, and product classification.

Risk Management
Model/Version Traceability
Intended-Use Documentation
Data Provenance
Human Oversight
Cybersecurity Considerations
Change Management
Monitoring

Validation shouldn't stop
when the model goes live.

1

BEFORE DEPLOYMENT

Validate

2

DEPLOYMENT

Observe

3

REAL WORLD

Monitor

4

MODEL UPDATE

Re-evaluate

5

CONTINUOUS IMPROVEMENT

Learn

Where we're starting

Research Phase · Prototype

Medical Imaging

Initial research focus: Radiology AI

First research question

"How can we create a reproducible, vendor-neutral methodology for independently evaluating a radiology AI system?"

Building on existing work

We are researching and learning from existing standards, frameworks, and initiatives. Chernyx is exploring how these pieces can be operationalized into an independent evidence framework.

FDA Regulatory-Science Research
ACR Assess-AI
RSNA Benchmarking
CLAIM Guidelines
FUTURE-AI Framework
ISO Risk Management
Peer-Reviewed Literature

We're looking for collaborators

Chernyx Benchmark is an R&D initiative. We are interested in collaborating with people who understand the problem from different sides.

Radiologists

Clinical perspective

Medical AI Researchers

Evaluation methodology

Biostatisticians

Statistical rigor

Healthcare & Regulatory Experts

Governance and implementation

Research Status

01

Research

Understanding the existing landscape

02

Methodology

Defining the benchmark

03

Prototype

Building the evaluation engine

04

Validation

Testing the methodology

05

Collaboration

Working with clinical and research partners

Medical AI doesn't just need better models.

It needs better evidence.

Chernyx Benchmark™ — A Chernyx AI Labs Research Initiative

Research phase. Methodology under development.