Skip to main content
Bright Minds. College Leslie Nichols
A student's gloved hand using a stainless probe to identify a structure on a preserved sheep brain on a black dissection tray; an instructor's clipboard with a grading rubric is visible in soft focus, with additional organ specimens (a heart, a kidney pair) on trays in the background.
Different instruments. Different evidence. Different claims.
Lab Notes · Front-line · Essay 03

The competency I can certify after a practical, and the one I can’t after a multiple-choice exam.

Disagreement about assessment in undergraduate science is often framed as disagreement about values: tradition vs. innovation, standards vs. equity, rigor vs. accessibility. A more useful starting question is what evidence an assessment provides about a student's knowledge or skill.

Leslie Nichols, M.S. Former lab coordinator & instructor · ~9 min read

Every assessment provides evidence for a claim about a student. A well-designed multiple-choice exam can show that the student can select a correct answer from plausible alternatives under the test's conditions. Depending on the questions, that may require factual knowledge, interpretation, or applied reasoning.

A lab practical provides different evidence. It can show that the student can identify relevant features on a specimen or demonstrate a procedure against stated criteria. The claim depends on which tasks are sampled and how they are scored; one successful station does not establish competence in every related task.

Neither instrument is inherently better. They measure different constructs: the knowledge or skills being assessed. The question for any program is not "which one is real assessment" but "which claim does this course need to be able to defend?" That question has an answer, and the answer is course-specific.

What a multiple-choice exam can and cannot certify

Multiple-choice exams can sample a broad range of content and are efficient to administer and score. Well-designed questions can help track learning over a term and prepare students for professional exams that use the format. Question quality and coverage still matter.

Multiple-choice questions can assess application through cases and problems. What they cannot directly show is whether a student can carry out an action with a specimen, instrument, or patient. Selecting a sound plan and performing it are related but different evidence.1 The AERA / APA / NCME Standards for Educational and Psychological Testing provide a framework for asking whether evidence supports the interpretation and use of a score.2

What a lab practical can and cannot certify

A lab practical is a performance assessment. A student works with a specimen or instrument and produces an observable response that can be scored against a defined criterion. It can assess identification, procedure, and decisions made during the task. Its results describe performance under those conditions, not readiness for unsupervised clinical practice.

A practical takes time, physical materials, supervision, and trained scorers. Medical education offers a useful comparison: the Objective Structured Clinical Examination (OSCE), which assesses clinical skills at a series of stations. Harden and Gleeson's 1979 paper describes this approach.3 Later research examines how task sampling, checklists, examples, and scorer training affect reliability. These features need evaluation; a practical is not reliable simply because it involves a real task.4

Figure 1 · Four assessment instruments by cost and construct
A 2x2 matrix placing four assessment instruments by cost to administer and construct measured. Cost to administer ↑ Construct measured → High Low Recognition Application Multiple- choice Short- answer Lab practical OSCE EXPENSIVE / RECOGNITION EXPENSIVE / APPLICATION CHEAP / RECOGNITION CHEAP / APPLICATION
This is a simplified comparison of typical uses and costs, not a fixed ranking. Multiple-choice questions can assess reasoning as well as recognition; observed tasks provide direct evidence of performance.

Norm-referenced vs. criterion-referenced: a real distinction

Two approaches to interpreting a score are worth distinguishing. A norm-referenced instrument tells us how a student performed relative to other students, using a percentile rank or a comparison with the class, for example. A criterion-referenced instrument tells us whether a student has demonstrated a defined competency, independent of how anyone else performed. A practical graded against published criteria is an example. Either approach can use written or performance-based assessments.

The distinction was formalized by Glaser in 1963 and has been a foundation of educational measurement ever since.5 Both frameworks are legitimate; they answer different questions. Mastery learning, in the educational-research sense Bloom developed, is criterion-referenced by definition: the standard does not move based on cohort performance, and students who do not yet meet it are given additional opportunities until they do.6 The Kulik meta-analysis found learning benefits across a range of mastery-learning programs, while also examining differences among settings and approaches.7

Assessment before professional training and licensure

When a course prepares students for professional training, assessment should show whether they have learned the knowledge and skills the next program expects. Published criteria make that expectation clearer than class rank alone.

Admission, certification, and licensure are different decisions. Admissions may rank applicants; licensing examines whether specified requirements have been met. A prerequisite course should prepare students for the relevant next step without claiming to certify the whole profession. Use observed performance when the learning goal is a procedure, and use written questions where they provide appropriate evidence.

A letter grade alone does not tell me which skills a student demonstrated. Published criteria and an observed performance give me a clearer basis for that judgment.

The work required for consistent scoring

The strongest argument against expanding the use of practicals is, fairly, a measurement argument: practicals are harder to score reliably than multiple-choice exams. This is a real challenge that a curriculum committee needs to address. Two raters watching the same performance can disagree on its quality; the same rater on a different day can score the same performance differently.

OSCE research offers practical ways to investigate and improve consistency: clear scoring criteria, shared examples, scorer training, and sufficient sampling of tasks.4 Before an assessment, graders can score the same examples and discuss disagreements. During and after it, the program can check whether the scores support the intended decisions.

What this means for grading philosophy

If a course's purpose is to assess competency, I favor clear, published standards and opportunities to improve after an unsuccessful attempt. A grade of 95% on a third attempt and 95% on a first attempt may reflect the same current performance when the criteria and conditions are comparable. A course should state separately how attempts affect its final grade.

The grading approach should match the decision being made. Ranking may serve a selection process, but rank alone does not establish that every student has met a competency standard. Courses can report achievement against that standard without treating another student's performance as the measure.

A practical implication

For a course preparing students for professional study, review each major assessment. Name the knowledge or skill it measures and the claim its scores support. Then ask whether that claim matches the course's learning goals. Retain useful assessments; revise those that leave important goals untested.

This review may lead to a different mix of written questions and observed tasks. The aim is not to favor a format. It is to gather enough relevant evidence to explain what the course's grades mean.

References & further reading

  1. Norman, G. R., Swanson, D. B., & Case, S. M. (1996). “Conceptual and methodological issues in studies comparing assessment formats.” Teaching and Learning in Medicine, 8(4), 208–216. doi:10.1080/10401339609539791. One of the canonical short references on what multiple-choice instruments can and cannot measure in health-professions education. See also van der Vleuten, C. P. M., & Schuwirth, L. W. T. (2005), “Assessing professional competence: from methods to programmes,” Medical Education, 39(3), 309–317.
  2. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. Washington, DC: AERA. testingstandards.net. The field's settled rulebook on validity, reliability, and fair use of assessment instruments.
  3. Harden, R. M., & Gleeson, F. A. (1979). “Assessment of clinical competence using an objective structured clinical examination (OSCE).” Medical Education, 13(1), 41–54. doi:10.1111/j.1365-2923.1979.tb00918.x. The originating paper for the OSCE as a structured performance assessment.
  4. Brannick, M. T., Erol-Korkmaz, H. T., & Prewett, M. (2011). “A systematic review of the reliability of objective structured clinical examination scores.” Medical Education, 45(12), 1181–1189. doi:10.1111/j.1365-2923.2011.04075.x. A systematic review of OSCE score reliability and the assessment-design factors associated with it.
  5. Glaser, R. (1963). “Instructional technology and the measurement of learning outcomes: some questions.” American Psychologist, 18(8), 519–521. doi:10.1037/h0049294. The originating paper on the criterion-referenced / norm-referenced distinction.
  6. Bloom, B. S. (1968). “Learning for mastery.” UCLA Evaluation Comment, 1(2), 1–12. The foundational treatment of mastery learning as criterion-referenced instruction with multiple opportunities to demonstrate.
  7. Kulik, C.-L. C., Kulik, J. A., & Bangert-Drowns, R. L. (1990). “Effectiveness of mastery learning programs: a meta-analysis.” Review of Educational Research, 60(2), 265–299. doi:10.3102/00346543060002265. The meta-analytic synthesis of mastery-learning outcomes against time-based instruction.

Drafted May 2026.