A Tool Can Generate a Prototype. It Can't Tell You If Your Analysis Was Actually Finished.
I've been watching people talk about Claude Design like it's about to make half the instructional design career field obsolete. Describe a course, get a prototype back in minutes — no Figma, no coding, no waiting on a developer. It's a genuinely useful tool, and I don't doubt the excitement is sincere. But "revolutionize the field" is doing a lot of work in that sentence, and it's worth actually unpacking what it would take for that to be true.
Here's what a tool like that can't do, no matter how good the underlying model gets: it can't tell you whether the analysis it's building from was actually finished.
That's not a small caveat. It's the entire job. ADDIE isn't a five-letter acronym you memorize and then execute in sequence — Analysis, Design, Development, Implementation, Evaluation, check the boxes, done. Applying it is a skill, and most of that skill lives in the Analysis phase, which is exactly the part that's invisible in a finished prototype. Nobody looking at a slick mockup can tell whether the interviews behind it actually surfaced the real performance gap, or just confirmed what the SME already assumed going in.
Take a concrete example. A stakeholder tells you employees "don't understand the safety procedure." A practitioner who actually knows how to apply ADDIE doesn't take that at face value — they ask what "don't understand" means operationally. Is it a knowledge gap, or is it that the procedure itself is unworkable on the floor and everyone's improvising around it? Those are two completely different training solutions, and one of them isn't a training solution at all — it's a process fix that no course will ever touch. A generative tool can build you a beautiful course for either interpretation. It has no way of knowing which one you actually need, because that judgment call happens in a conversation, in a follow-up question, in noticing that the stakeholder's frustration doesn't quite match what the frontline workers are describing. None of that produces a prompt. It produces a conclusion a practitioner reaches by actually doing the analysis.
Or take Kirkpatrick's Level 3 and Level 4 — behavior change and business results, the two levels most people entering this field skip past because they're hard to measure and harder to attribute. Knowing how to design toward those levels, instead of stopping at "did they pass the quiz," is a skill you develop by watching training fail to change behavior and figuring out why. A tool can generate an assessment. It can't tell you that the assessment is measuring recall instead of the thing that actually matters, because it doesn't know what mattered in the first place — it only knows what you told it, and if what you told it was incomplete, the output will be a polished version of that same incompleteness.
A tool also can't be in the room.
I ran a pilot once for a program where we'd completely restructured the training — pulled all the knowledge-based content out of the ILT and moved it into an eLearning course, then rewrote the ILT itself as a single day of real-world exercises. We built those exercises from scratch, designed to simulate the actual work. On paper, it was a clean split: eLearning for knowledge, ILT for application. In the room, it fell apart by lunch. The learners didn't have enough context to understand what the exercises were actually building toward — we'd assumed the eLearning had given them enough of a frame to walk into hands-on work cold, and it hadn't. You could see it happening in real time: the confused looks, the exercises stalling out, people going through the motions without understanding why. We stopped the pilot at lunch and went back to rebuild the exercises using actual artifacts and items from the workplace instead of the ones we'd invented, so the exercises finally connected to something learners already recognized.
No prototype would have caught that. A generated course can look complete, sound complete, hit every objective on the outline — and still fail the moment real people are sitting in a room trying to do the thing you designed. The gap between "this should work" and "this is working" only shows up in front of an actual audience, watching actual faces, in real time. That's not a data point you feed into a tool afterward. It's a judgment call you make on the spot, mid-pilot, deciding the exercises are broken before the day is even over — and then doing the harder work of rebuilding them from something real.
This is why I don't think the field is being revolutionized so much as it's being handed a much faster front end. The prototype used to take days to mock up. Now it takes minutes. That's real, and it's useful, and I'll take it. But the thing that separates a course that works from a course that just looks finished was never the mockup stage. It was whether someone did the unglamorous work of figuring out what was actually broken before anyone touched a storyboard, and whether someone was in the room to catch it when the design didn't hold up against real people. A tool can't do either of those things. It can only build, very quickly and very convincingly, on top of whatever you hand it — including your mistakes.
That's the part of the job that doesn't get revolutionized by a better prompt box. It gets learned, slowly, by doing the analysis wrong a few times, watching a pilot fall apart in real time, and paying attention to what that cost.