Commit enough studies to memory and the marks will follow. That’s the working assumption for a lot of IB Psychology SL students, and it’s the wrong variable to optimize. The course is research-heavy, which makes volume feel like the lever to pull—but examiners aren’t measuring how much a student has memorized. They’re measuring what a student does with what they know.
That logic runs through every assessment component, from short-answer questions to extended-response tasks. The skill tested consistently across every question format is using evidence to answer a specific question—not demonstrating that the evidence was stored.
Table of Contents
Studies as Tools, Not Facts
The three core approaches—biological, cognitive, and sociocultural—are examined through short-answer and extended-response formats. The SL elective is assessed through Paper 2 extended-response questions carrying the same command terms and the same demand to build a supported argument. What earns marks in both is argument: deploying named findings to answer a specific question, with each study doing evidential work. Moving through every approach, cataloging what researchers found, can produce an accurate account of the field—but accurate description, however thorough, doesn’t satisfy an evaluative command term.
Practitioner guidance from Eye Opener (ibpsychology.org, 2026) puts it directly: when students understand how psychological research is produced and what role they play in using it, the course shifts from memorizing isolated studies to disciplined critical thinking. Exam responses become “reasoned rather than rehearsed” because studies function as evidence in an argument rather than stand-alone facts. In extended-response questions, students synthesize established evaluation—recognized methodological critiques, replications, and published debates—to reach a question-bounded conclusion, rather than re-judging original research from scratch. In the Extended Essay, they again work with secondary data, using abductive reasoning to move between theory and evidence, adding their own reasoned judgments by comparing findings across theoretical, methodological, and cultural contexts.
There’s a distinction worth drawing here. Description is an accurate account of what a study did and found—enough to make it usable in an argument. Evaluation in ERQs is something different: a question-linked judgment that weighs how much a study’s findings support the specific claim, draws on recognized methodological issues or established critique, and arrives at a conclusion about what the evidence can and cannot justify. Generic limitation-listing—naming flaws with no connection to the claim being argued—isn’t evaluation; it’s the appearance of critical thinking without the work. In essays, students’ authority comes from synthesizing established evaluation and forming reasoned, bounded judgments. In the Internal Assessment (IA), that authority comes from reflexively evaluating their own methodological and ethical decisions.

The Internal Assessment—A Related but Distinct Form of Evaluation
The IA develops evaluative skill in a different register. Students act as junior quantitative researchers running a partial replication of an existing study, and the evaluative task is reflexive rather than external: they must account for how their own methodological choices—sampling decisions, operationalizations, and controls—shaped the data and the conclusions they can draw. Ethical reflexivity is a distinct part of this, as students consider how well they obtained informed consent, honored participants’ right to withdraw, minimized stress or deception, protected anonymity and confidentiality, and adhered to IB ethical guidelines. The IA isn’t appraising published research; it’s treating students’ own methods as decisions with consequences.
Eye Opener draws the definitional line cleanly. Reflexivity is an internally driven appraisal of one’s own research choices; evaluation is an externally directed judgment of another study’s validity and credibility. In ERQs, evaluation means synthesizing established critique rather than re-adjudicating studies independently. In the IA, the demand turns inward—toward the student’s own replication, procedural choices, and the conclusions those choices can and cannot justify. In every component, the type of evaluation expected is bounded: in the IA by research role, in the exam by the specific cognitive operation the command term names.
Command Terms as the Real Assessment Differentiator
The difference between “describe,” “discuss,” “evaluate,” and “to what extent” is not cosmetic—each names a different cognitive operation. “Describe” calls for accurate reproduction of what a study did and found. “Discuss” requires a balanced review that weighs competing arguments. “Evaluate” demands explicit appraisal of strengths and limitations. “To what extent” requires a defended, evidence-backed position. IB subject guidance defines these as the fixed reference points students need before practicing any response format. Misreading the verb—delivering description where evaluation is required, or a one-sided argument where balance is expected—is a reliable source of mark loss even when content knowledge is strong.
The difference becomes concrete when the same study is used twice. In AO1 mode: method, key finding, and basic link to theory—accurate, but reproductive. In AO3 mode, that same study becomes evidence weighed for the question: the finding connects to an answer claim through explicit reasoning, a limitation is named and linked to what the study can and cannot conclude for this specific question, and the response closes with a bounded conclusion about what the evidence can and cannot justify. According to the IB’s official 2026 examiner instructions, Paper 1 Section B questions are worth 22 marks and all carry an AO3 command term—a response in AO1 mode is structurally capped regardless of how many studies it includes.
A 2021 peer-reviewed study accessed via the ERIC database found substantial variation in how nine experienced teacher-assessors and six published sources defined command words including “evaluate” and “analyse”—including fundamental disagreements about whether certain terms require a judgment at all—and concluded that this inconsistency risks confusing students and weakening assessment reliability. Relying on intuition rather than internalizing the official definitions is an avoidable preparation risk.
Building the Skill Across Two Years
The ability to deploy studies as evaluative evidence develops in sequence, not all at once. Accurate study summaries come first: students need a precise understanding of what a study found and how it was conducted before any evaluative use of it is possible. The intermediate step is practicing command-term-specific responses with the same study—writing a description, then a discussion, then a full evaluation. That repetition turns an abstract distinction into a felt one.
Timed extended-response rehearsal follows once a student can hold content and argument structure simultaneously. Platforms that organize practice by command term and provide structured feedback can help compress this cycle—Revision Village is one widely used example—but no platform substitutes for the prior habit: read the command term, understand what it demands, then write.
Thinking With Studies Across IB Psychology SL
IB Psychology SL is structured around a specific evaluative skill—not content storage, but the ongoing demand to use evidence to answer a question. That demand runs through every component. Students who internalize it early can read a command term and know exactly what kind of argument is required—which narrows down what each named study needs to do and what conclusion the evidence can actually support.