Pilot Simulation and Data-Informed Evaluation
A formative evaluation reflection examining what worked, what created friction, and how data from a virtual pilot session with high school students directly shapes the next design iteration.
Once a learning solution is delivered, how do you know it actually worked? That question sits at the center of this reflection. Rather than treating delivery as the finish line, this evaluation examines the pilot session as a source of data — a first real test of the design's assumptions about learners, sequence, and cognitive demand.
The implementation context
The pilot simulation was reviewed on April 24, 2026 and ran approximately 45–50 minutes in a virtual asynchronous format. The target learner population was high school students — a group whose prior knowledge, attention patterns, and technical access vary considerably. That variability is itself a design constraint: materials designed for a general high school audience must work for learners across a wide range of backgrounds without assuming shared baseline knowledge.
During the session, I served as both designer and evaluator simultaneously — observing how learners moved through the experience while actively documenting observations about clarity of directions, transitions between activities, pacing across the session, and overall cognitive demand. That dual role requires discipline: the instinct to intervene or explain must be suppressed so that the design itself can be evaluated, not the facilitator's ability to compensate for design gaps.
The hardest part of evaluation is watching something not work — and writing it down instead of fixing it in the moment.
What worked well
Objective-activity-assessment alignment. The strongest element of the design was the alignment between stated objectives, learning activities, and assessment tasks. Each activity connected directly to a measurable outcome, which meant learners could complete the session without wondering why any given task was included. That alignment is a fundamental design principle (Anderson & Krathwohl, 2001), and observing it hold up in a pilot context provided meaningful confirmation.
Instructional flow. The sequence — self-assessment → exploration → reflection → application — created a logical progression that built on itself. Learners weren't asked to apply knowledge they hadn't yet constructed, and the transition from reflection to application gave the final activities a sense of purpose rather than appearing as disconnected tasks.
Higher-order thinking activities. The reflection and comparison activities pushed learners beyond recall into genuine analysis and evaluation — the upper levels of Bloom's revised taxonomy (Anderson & Krathwohl, 2001). Asking students to justify a career choice based on their own data rather than select a preset answer creates a qualitatively different kind of learning. The pilot confirmed that these activities were engaging and that learners were willing to do the harder thinking when the prompt was clear.
What needs improvement
Clarity of instructions. Some written instructions — particularly around the career comparison activity — required learners to read through the directions more than once before proceeding. Rereading is a reliable signal that either the instruction is ambiguous, the task structure is unclear, or the example provided isn't representative enough. This was the most consistent friction point across the session (Clark & Mayer, 2016).
Cognitive load in multi-criteria tasks. The career comparison worksheet asks learners to evaluate multiple careers across multiple criteria simultaneously. While the parallel structure is logical from a design standpoint, presenting all criteria at once creates a working memory demand that likely exceeds what's necessary for the comparison objective (Sweller et al., 2019). The result is that some learners spent more cognitive effort navigating the format than doing the comparison itself.
Pacing and chunking. A 45–50 minute virtual session is long for an asynchronous format without built-in breaks or transition signals. The pilot revealed that the later activities — which require the most synthesis — were reached when learner attention was most taxed. This is a pacing and sequencing problem, not a content problem.
Formative evaluation questions
The following four questions drove the evaluation and would be used to structure learner feedback collection in a live implementation:
- Which parts of the session were confusing or unclear?
- Were there any points where you felt like pausing or stopping?
- Which activities felt most helpful for your thinking?
- How confident do you feel about your career direction now, compared to when you started?
These questions are designed to surface both instructional quality (questions 1 and 2) and learning effectiveness (questions 3 and 4) — aligning with Kirkpatrick's model for separating learner reaction from actual learning outcomes (Kirkpatrick & Kirkpatrick, 2006).
Future modifications
The evaluation data points to three concrete revisions:
- Simplify and exemplify directions. Every set of activity instructions will be rewritten with a worked example included. The example should demonstrate the exact format expected, not just describe it abstractly — following Clark and Mayer's (2016) guidance on reducing extraneous load through concrete models.
- Break multi-criteria tasks into smaller sections with clearer transitions. The career comparison worksheet will be restructured so learners complete one career's evaluation before moving to the next, rather than filling out a full parallel grid. This reduces working memory demands and makes the task feel more manageable (Sweller et al., 2019).
- Add explicit transition signals and chunking breaks. The session will be restructured into three clearly labeled segments with brief transition prompts between them, giving learners a sense of progress and a moment to consolidate before the next section begins (Branch, 2018).
Revision isn't failure — it's the evaluation working as intended. A design that generates no revision data wasn't tested rigorously enough.
References
- Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A taxonomy for learning, teaching, and assessing: A revision of Bloom's taxonomy of educational objectives. Longman.
- Black, P., & Wiliam, D. (2009). Developing the theory of formative assessment. Educational Assessment, Evaluation and Accountability, 21(1), 5–31.
- Branch, R. M. (2018). Instructional design: The ADDIE approach. Springer.
- Clark, R. C., & Mayer, R. E. (2016). E-learning and the science of instruction (4th ed.). Wiley.
- Gustafson, K. L., & Branch, R. M. (2007). What is instructional design? In R. A. Reiser & J. V. Dempsey (Eds.), Trends and issues in instructional design and technology (2nd ed., pp. 10–16). Pearson.
- Hmelo-Silver, C. E., Duncan, R. G., & Chinn, C. A. (2007). Scaffolding and achievement in problem-based and inquiry learning. Educational Psychologist, 42(2), 99–107.
- Kirkpatrick, D. L., & Kirkpatrick, J. D. (2006). Evaluating training programs: The four levels (3rd ed.). Berrett-Koehler.
- Molenda, M. (2015). In search of the elusive ADDIE model. Performance Improvement, 54(2), 40–42.
- Sweller, J., van Merriënboer, J. J. G., & Paas, F. (2019). Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 31(2), 261–292.