Page Menu

Independent Assessment of AI-Enabled Continuing Medical Education

UMass Chan evaluates learner-facing AI systems that generate clinical content for continuing medical education (CME). Our version-specific assessments use predefined clinical scenarios, expected-response guides, standardized scoring, documented quality controls, and transparent reporting to identify strengths, limitations, and potentially consequential errors.

Who this service is for? 

This service is designed for companies and organizations developing or operating AI-enabled educational activities for physicians and other health professionals.

  • Learner-facing clinical question-answering or search systems
  • AI-generated clinical explanations, recommendations, or summaries
  • Tools that retrieve, synthesize, or cite clinical evidence for educational use
  • Existing activities being updated after a material change to the model, configuration, evidence source, or learner workflow

An initial consultation determines whether the activity and requested scope are suitable for this service.

What we assess? 

We assess a defined activity under defined conditions. We assess the submitted educational activity, not the company or an abstract model as a whole. Each assessment is tied to the intended learners and use, clinical scope, learner workflow, system or model version, configuration, retrieval or citation settings, testing environment, and evaluation period.

The assessment may include the system's responses to predefined clinical scenarios, the evidence displayed to the learner, warnings and limitations, consistency across repeated or varied prompts, and performance within the documented workflow.

The findings apply only to the documented activity and testing conditions. A material change may require limited or full reassessment.

What we evaluate? 

Clinical accuracy
Are clinical statements, interpretations, recommendations, and cited evidence consistent with accepted evidence and practice?

Completeness and context
Does the response include the limitations, alternatives, contraindications, uncertainty, and context needed for appropriate interpretation and use?

Potentially harmful or misleading content
Could an error, omission, or framing problem reasonably contribute to misunderstanding, inappropriate care, or patient harm?

Balanced treatment of options
Does the response address reasonable alternatives and avoid unsupported preference, promotional framing, or other commercial influence?

Evidence and citation traceability
Are cited sources verifiable, relevant, sufficiently current, and supportive of the claims made?

Consistency and limitations
Does performance remain reasonably consistent across repeated or varied prompts, and are uncertainty and scope limitations communicated clearly?


How the assessment works? 

  1. Initial consultation and fit review

    We learn about the intended learners, clinical purpose, activity workflow, and current development status and determine whether the proposed work is a fit.

  2. Scope and readiness

    We document the system or model version, configuration, access, content sources, clinical scope, testing environment, data-protection conditions, and evaluation period.

  3. Evaluation set and rubric

    UMass Chan develops and controls a set of predefined clinical scenarios, expected-response guides, and scoring criteria appropriate to the intended use. Company input may inform context and representative scenarios but does not control the final evaluation criteria or findings.

  4. Controlled testing and evidence capture

    Approved scenarios are run through the documented learner-facing workflow. We capture prompts, outputs, visible sources, warnings, timestamps, and relevant system and configuration conditions.

  5. Standardized scoring and quality assurance

    Responses are evaluated against expected-response guides using a documented rubric. Quality-control checks are applied, and ambiguous, disputed, or high-risk findings are escalated to appropriate subject-matter experts. Automated methods may support structured analysis but do not independently determine the significance of high-risk findings.

  6. Reporting and reassessment planning

    UMass Chan provides a structured report that describes the test conditions, findings, limitations, and appropriate-use considerations and identifies changes that may trigger limited or full reassessment.

Consequential failure modes 

Designed to surface consequential failure modes
Depending on the intended use and clinical scope, the evaluation set may examine:

  • Medication selection, dosing, contraindications, and interactions
  • Special populations and urgent escalation
  • Hallucinated, unsupported, or outdated evidence
  • Unwarranted certainty, flawed premises, and omitted uncertainty
  • Fringe or disproven practices
  • Omission of reasonable diagnostic or therapeutic alternatives
  • Unsupported product preference or promotional framing
  • Inconsistent answers across repeated queries or changes in prompt wording

Coverage is tailored to the activity's intended use. The controlled evaluation set is finalized before testing and kept confidential to preserve test integrity.

What the organization receives? 

  • A record of the activity, system version, configuration, learner workflow, and evaluation period
  • A summary of the assessment method, scenario categories, and scoring approach
  • Quantitative and qualitative findings across the evaluation domains
  • Itemized material discrepancies with supporting rationale and evidence
  • Documented strengths, limitations, and appropriate-use considerations
  • Recommended remediation priorities and triggers for reassessment

Detailed evidence, including complete prompts, outputs, expected-response guides, scoring records, and reconciliation documentation, may be made available to authorized parties under an applicable agreement. Controlled test materials remain protected.

Independence and data protection 

How we protect the independence of the assessment

  • UMass Chan controls the evaluation methodology, final scenario set, expected-response guides, scoring criteria, analysis, and report.
  • The company provides the access and factual information needed to test the submitted activity and may provide representative scenarios, but it does not determine the final criteria or findings.
  • Financial terms are agreed in advance and are not contingent on favorable findings, remediation recommendations, or commercial outcomes.
  • Favorable and unfavorable findings are documented. Relevant organizational and reviewer conflicts are identified and managed.
  • The assessment is an independent report of findings, not a product endorsement or promotional seal.

Protecting confidential information and test integrity
Assessment materials are handled in UMass Chan-authorized systems with role-based access and institutional controls. Controlled scenarios, expected-response guides, and scoring anchors are access-restricted and version-controlled.

Submitted data, prompts, system outputs, and assessment records are not used to train or improve an AI model or external AI service without separate written authorization and required institutional review.

Do not send protected health information, learner data, credentials, proprietary test items, or confidential system output by email. A secure intake path will be provided when appropriate.

Important limitations and leadership 

Important limitations
The assessment provides independent findings for the defined activity, system version, configuration, and evaluation period. It does not confer CME credit, determine eligibility for CME credit, certify or endorse a product, constitute regulatory approval, or guarantee future performance. The assessment does not evaluate an entire platform, source code, cybersecurity posture, clinical outcomes, or every possible learner question unless expressly included in the agreed scope.

UMass Chan leadership and expertise
Assessments are led by the UMass Chan AI Assurance Lab. The evaluation team draws on expertise in clinical informatics, biostatistics, AI evaluation, continuing education, implementation science, data governance, and quality assurance. Team composition is matched to the activity's clinical scope and risk. View the AI Assurance Lab leadership

Request an initial consultation 

To help us determine fit and plan the discussion, please be prepared to describe:

  • The activity, intended learners, and educational use
  • The clinical scope and learner workflow
  • The current system or model version and configuration
  • The access available for controlled testing
  • Known limitations and principal content or evidence sources
  • Funding, sponsorship, and relevant commercial relationships
  • The requested timing and current stage of development

Request an initial consultation

Do not send protected health information, learner data, credentials, proprietary test items, or confidential system output.