How we grade
The whole ruler, before you start. Nothing here changes after the score is out.
How the score is calculated
Every decision point has an AI-generated answer key, with a cited source, and it can be wrong. In multiple choice, the sum is a fixed rule: the same choices give the same score. In written answers, a language model judges what you wrote and can also be wrong.
Safety is a ceiling, not a weight
Three axes enter the average — decision, reasoning and context. The fourth, safety, does not: a critical error caps the final score from above, no matter how good the rest is.
A critical error was found in one decision in this case, so the final score is capped below the competency average — safety works as a ceiling, not part of the average.
Affected: decision 3.
One critical error in one decision. This is the actual design, and it's the same number you see in the verdict.
Where the ruler comes from
- Brazilian guidelines firstSBC, SBP, Febrasgo and similar bodies, with recommendation class and evidence level when they exist.
- Every point cites its sourceYou read the citation and the literal excerpt inside the verdict, not in a bibliography in the footer.
- No guideline, we say soA point that no guideline grades is marked as such. That's not a missing source: it's a point about something else.
If you disagree
Every graded decision has a dispute button. The dispute is recorded in a team review queue (it is not a medical review). The review happens outside the product: no deadline and no notice.
The four axes that make up the score
- DecisionThe right call at the decision point
- ReasoningQuality of the clinical justification
- ContextUse of patient state in the decision
- SafetyNot part of the average — one critical error caps the final score, no matter how strong the rest is.
Educational tool. Nothing here is guidance for a real patient, and no score from this product certifies clinical competence (CFM Resolution 2,454/2026).