← All posts
AI in education licensure exam prep NCLEX adaptive learning

The Exam Already Agrees With You

Patrick Sullivan, Ph.D. ·
Also posted on
During the NCLEX-RN licensure exam, the study materials stay outside. What the exam measures is the clinical judgment the learner has built. This image was generated by Patrick Sullivan through an iterative process with OpenAI’s GPT.
During the NCLEX-RN licensure exam, the study materials stay outside. What the exam measures is the clinical judgment the learner has built. This image was generated by Patrick Sullivan through an iterative process with OpenAI’s GPT.

When a proctored, adaptive exam is the endpoint, borrowed fluency has nowhere to hide and productive struggle has to be budgeted, not just preserved.

Nursing licensure is where the learning-environment argument stops being a philosophy and starts being measured by an instrument built to detect exactly the difference between borrowed fluency and built understanding.

After reading William’s post about AI and the learning environment (AI Is Not a Content Problem. It Is a Learning Environment Problem. Aug. 4th), I want to reinforce some of his ideas and challenge others. I have spent my career in academic research (environmental health, land-use modeling, human and community development) settings where the finished artifact — the paper, the report, the model run — was always taken as evidence of the thinking behind it. For the past several months I have been working with the RokSpark team on a different problem: helping people prepare for professional licensure examinations, beginning with the NCLEX-RN, the National Council Licensure Examination that every registered nurse in the United States must pass to practice.

In his Aug. 4th post, William argued that AI in education is not primarily a content problem but a learning-environment problem: what matters is not the answers a system can produce but the understanding a learner builds along the way. I want to press on that argument from the domain I have been working in, because licensure preparation does two things to William’s thesis at once. It confirms it more decisively than any setting I know. And it adds constraints his essay does not address.

An exam built to detect mental models

William’s central worry is that a polished product no longer tells us whether its author understands anything. It may represent sustained thought, or merely a well-phrased prompt and thirty seconds of waiting. In most of education, that is a crisis of inference: the artifact was our evidence, and now the evidence has been contaminated.

Nursing licensure solved this problem by refusing to accept artifacts at all. The NCLEX is administered in a proctored testing center, adapts its difficulty to the candidate’s responses in real time, and produces nothing that can be drafted, revised, or delegated. A candidate cannot prompt their way through it. Whatever fluency they bring into that room is fluency they built.

But the more interesting fact is what the exam now measures. In April 2023, the National Council of State Boards of Nursing (NCSBN) launched the Next Generation NCLEX, rebuilt around its Clinical Judgment Measurement Model: a framework developed because the evidence showed that newly licensed nurses were being asked to make complex care decisions that knowledge-recall testing did not predict. The exam now works through unfolding case studies that follow a six-part cognitive sequence: recognize cues, analyze them, prioritize hypotheses, generate solutions, take action, evaluate outcomes.

Set that sequence beside the definition of a mental model William inherited from Rachel Kaplan — a simplified internal representation that lets us recognize patterns, anticipate what might happen next, weigh alternatives, and act. The correspondence is nearly term for term. NCSBN’s psychometricians, working from decision science and clinical practice analysis, converged on the same account of understanding that the environmental psychologists reached from studying how people make sense of their surroundings. The highest-stakes examination in American healthcare is, in substance, a mental-model detector.

So the licensure domain does not merely illustrate William’s thesis. It supplies the external referee his argument otherwise lacks. If a learning environment produces convincing exteriors rather than built understanding, the exam will say so — publicly, consequentially, and in a form no one can argue with.

The learners with the least attention to spare

William wrote that directed attention is limited and exhaustible, and that a learning environment can spend it wisely or squander it. Consider who is actually preparing for the NCLEX: a student in the final semester of a nursing program, or a recent graduate, often juggling clinical rotations, a job, and a family, frequently holding an employment offer contingent on passing, with a test date fixed on the calendar. There may be no population of adult learners whose directed attention is under heavier tax.

Now consider what the preparation market hands them. The dominant products are, in William’s image, dump trucks: question banks of several thousand items, libraries of generic video lectures, and the standing instruction to do a hundred questions a day. The diagnostic work — figuring out what you don’t know — is left to the candidate. So is the scheduling: deciding what to study tonight, three weeks out, tired. So is progress management: judging whether you are ready. The incumbent model does not squander attention incidentally. It squanders attention structurally, by offloading exactly the cognitive burdens a learning environment exists to carry.

The consequences are not evenly distributed. In 2025, first-time U.S.-educated candidates passed at 86.7 percent, but repeat candidates passed at barely more than half that rate, and the all-candidate figure fell to 69.1 percent. Every failed attempt means a mandatory waiting period, another examination fee, delayed income, and — this is the part the statistics do not show — a blow to confidence that makes the next attempt harder. When roughly one candidate in three, across the whole pool, walks out of the testing center without a license, the quality of the preparation environment is not a consumer-choice question. It is a workforce question and an equity question, arriving in the middle of a nursing shortage.

A learning environment with a deadline

Here is where licensure complicates William’s argument. His essay treats productive struggle as something to be preserved: resist answering too quickly, keep the learner in the territory between effortless completion and overwhelming frustration. I agree, and I would add a constraint he never has to confront. In licensure preparation, the learning environment has a deadline.

A candidate with an exam in twenty-one days cannot afford productive struggle wherever it happens to arise. Difficulty must be budgeted, not merely calibrated. The exam’s published test plan tells us which content areas carry the most weight; the candidate’s own performance tells us where they are weakest; the calendar tells us how much struggle there is time for. A Socratic detour on a low-weight topic the candidate has nearly mastered is not patient teaching. It is a misallocation of the scarcest resource in the room.

This means the environment itself must carry the burdens the incumbents offload: diagnose the starting point before instruction begins, anchor the plan to the test date, and redirect effort as the evidence about the learner accumulates. William asked what Lux needs to understand about a learner before responding. In licensure preparation, part of the answer is blunt: it needs to know what day it is.

What we can honestly claim

One more thing the licensure domain enforces, which I want to model here because the preparation market conspicuously does not: honesty about effect sizes.

The inspiration for much of this work is Bloom’s 1984 finding that students taught one-on-one outperformed classroom students by two standard deviations. That figure has not survived replication at that magnitude, and we should stop implying it will. The careful summary is VanLehn’s 2011 review, which put the effect of intelligent tutoring systems at roughly three-quarters of a standard deviation (nearly indistinguishable, in his analysis, from human tutors). That is a large effect by any educational standard. It is not a miracle, and the evidence on formative assessment more broadly ranges from strongly positive to genuinely contested.

In a market saturated with pass guarantees and testimonial marketing, stating the evidence at its actual size is both an ethical obligation and, I suspect, a durable advantage because licensure is the one domain where claims get audited. The endpoint is a pass-fail event, scored by someone else, on an instrument designed to detect whether understanding was built. If the learning-environment thesis is right, this is where it will show. And if we are wrong about any of it, this is where we will find out.

That is why I think licensure preparation is the proving ground for everything this publication is about: learners whose attention must be spent wisely, a deadline that forces design honesty, and a referee that cannot be fooled.

A question for readers: If you have prepared for a licensure or board examination — nursing, law, medicine, engineering, teaching — what did your preparation materials get right about how you actually learn, and what did they leave entirely on your shoulders?

The Learning Environment publishes a Monday essay, and a Thursday field note about what advances in AI mean for how humans learn. Subscribe to join the inquiry.