How we grade

The whole ruler, before you start. Nothing here changes after the score is out.

How the score is calculated

Every decision point has an AI-generated answer key, with a cited source, and it can be wrong. In multiple choice, the sum is a fixed rule: the same choices give the same score. In written answers, a language model judges what you wrote and can also be wrong.

Safety is a ceiling, not a weight

Three axes enter the average — decision, reasoning and context. The fourth, safety, does not: a critical error caps the final score from above, no matter how good the rest is.

Competency average82
Final score60

A critical error was found in one decision in this case, so the final score is capped below the competency average — safety works as a ceiling, not part of the average.

Affected: decision 3.

One critical error in one decision. This is the actual design, and it's the same number you see in the verdict.

Where the ruler comes from

  • Brazilian guidelines firstSBC, SBP, Febrasgo and similar bodies, with recommendation class and evidence level when they exist.
  • Every point cites its sourceYou read the citation and the literal excerpt inside the verdict, not in a bibliography in the footer.
  • No guideline, we say soA point that no guideline grades is marked as such. That's not a missing source: it's a point about something else.

If you disagree

Every graded decision has a dispute button. The dispute is recorded in a team review queue (it is not a medical review). The review happens outside the product: no deadline and no notice.

The four axes that make up the score

  • DecisionThe right call at the decision point
  • ReasoningQuality of the clinical justification
  • ContextUse of patient state in the decision
  • SafetyNot part of the average — one critical error caps the final score, no matter how strong the rest is.

Educational tool. Nothing here is guidance for a real patient, and no score from this product certifies clinical competence (CFM Resolution 2,454/2026).