Skip to content

How L&D Leaders Should Evaluate AI Use in Instructional Design

 

A lot of L&D leaders are now facing the same practical question: How do we know whether our instructional design team is using GenAI well?

That is not a small question.

In many organizations, AI use has already moved beyond curiosity. Teams are experimenting. Prompt libraries are circulating. Drafting speed is improving in places. Individuals are finding their own ways to use the tools. And leaders are being asked, implicitly or explicitly, whether this is working.

Unfortunately, many teams are still evaluating the answer too narrowly.

They look at:

  • adoption rates
  • number of active users
  • frequency of use
  • time saved on drafting
  • number of AI-assisted outputs produced

These things are not irrelevant. But they are not enough. In fact, on their own, they can be misleading.

Because an instructional design team can be using AI heavily and still not be using it well. It can be saving time while weakening review. It can be producing faster outputs while generating more polished mediocrity. It can look more productive while becoming more uneven in quality. It can show high usage without showing stronger instructional judgment.

That is why I believe L&D leaders need a better evaluation frame. The goal is not simply to know whether the team is using AI. The goal is to know whether AI is improving the team’s design capability, judgment, consistency, and performance.

That is a much stronger standard.

This article is part of a series on the future of instructional design in the age of GenAI. The series explores how instructional designers can move beyond ai hoc prompting toward a more disciplined, challenge-based human–AI working method.

Table Of Content

The Wrong Starting Point: Usage as the Main Measure

Let us start with the common trap.

When new tools are introduced, leaders often measure what is easiest to see. With GenAI, that tends to be usage.

  • How many people are using it?
  • How often?
  • For what tasks?
  • How much time is being saved?
  • How many deliverables are now AI-assisted?

These questions make sense at an early operational level. They tell you whether the tool is being adopted at all. That matters.

But usage is a very poor proxy for quality.

A junior designer may be using AI constantly because they lack confidence without it.
A stronger designer may be using it more selectively and more effectively.
One team member may be generating large amounts of polished weak content.Another may be using AI rarely but intelligently for critique and review.

If you evaluate success mainly through activity, you risk rewarding the wrong behavior. That is the first thing leaders need to correct.

High usage does not necessarily mean high maturity. Low usage does not necessarily mean resistance or weakness. And time saved does not automatically mean value created.

Those distinctions matter.

What Leaders Should Actually Be Evaluating

In my view, there are five better questions L&D leaders should ask when evaluating AI use in instructional design teams.

1. Is the quality of instructional design actually improving?

This is the first and most important question.

Not: Are people using AI?
But: Is the work better?

That means looking for signs such as:

  • stronger learning objectives
  • better alignment between objectives and assessments
  • fewer text-heavy screens
  • more thoughtful visual and interaction treatment
  • higher-quality formative assessments
  • clearer narration
  • fewer superficially polished but instructionally weak outputs

This is not always easy to measure in a dashboard. But it is still the right question.

Because if AI use is growing and design quality is not improving, then the team is not yet using AI well enough.

2. Is human judgment being strengthened or bypassed?

This is the question many leaders fail to ask.

It is possible for AI to help teams move faster while quietly weakening the quality of human thinking in the process. If designers start relying too heavily on AI structure, AI language, and AI-generated answers, they may become more efficient but less deliberate.

That is a serious risk.

So leaders should look for evidence of whether designers are still:

  • reviewing actively
  • challenging AI output
  • justifying choices
  • spotting weak logic
  • making real decisions rather than lightly editing generated content

If the human role is becoming thinner, that is a warning sign, even if the throughput looks better.

3. Is the team becoming more consistent?

One of the real opportunities of AI is improved consistency across a team. But this does not happen automatically.

In some teams, AI actually increases inconsistency because everyone uses it differently.

One designer uses it only for summarization.
Another uses it for nearly everything.
One reviews critically.
Another accepts too quickly.
One uses it for challenge and audit.
Another only for generation.

That kind of unevenness reduces the value of the tool at the team level.

So leaders should ask:

  • Are outputs becoming more consistently strong across IDs?
  • Are junior designers improving faster?
  • Is the gap between stronger and weaker designers narrowing in healthy ways?
  • Are teams following a more coherent method, or just using the same tool differently?

Consistency is one of the clearest signs that AI use is becoming operationally mature.

4. Is rework going down for the right reasons?

Rework is an important indicator, but it needs interpretation.

If AI is helping the team reduce avoidable drafting effort and improve first-pass quality, that is good. If fewer issues are being caught in review because reviews have become shallower, that is not good.

So leaders should not just ask whether rework is down. They should ask why.

For example:

  • Are SMEs asking for fewer structural changes?
  • Are reviewers identifying fewer alignment issues?
  • Are storyboards becoming cleaner earlier?
  • Are IDs catching weaknesses before formal review?
  • Or are people simply moving faster with less scrutiny?

This distinction is critical.

Less rework is valuable only if it reflects better thinking upstream, not weaker review downstream.

5. Is the team’s way of working becoming more disciplined?

This may be the most strategic question of all.

The long-term value of AI will not come simply from individual cleverness. It will come from stronger team method.

So leaders should evaluate whether the team is developing:

In other words, leaders should be asking not just, “Are people getting good at the tool?” but also, “Are we building a stronger operating model around the tool?”

That is a much more important question.

What Good Evaluation Might Look Like in Practice

A strong evaluation approach should combine both operational metrics and professional indicators.

Operational indicators may include:

  • time saved in specific stages
  • turnaround time
  • volume handled
  • cycle time from SME input to storyboard
  • number of review rounds

These are useful, but only as part of the picture.

Professional indicators should include:

  • quality of learning objectives
  • assessment quality
  • alignment strength
  • quality of review comments
  • reduction in text-heavy treatment
  • improvement in interaction relevance
  • level of independent reasoning shown by IDs
  • ability to critique AI output effectively

This second set is harder to quantify, but much more meaningful.

Leaders do not need perfect measurement to evaluate these things. They can sample projects, review patterns across deliverables, compare pre- and post-AI outputs, observe how IDs explain their choices, and speak with reviewers and SMEs.

The important thing is not to reduce the evaluation question to usage analytics alone.

What Leaders Should Watch Out For

There are a few warning signs that AI use may be heading in the wrong direction.

For example:

  • outputs are faster, but not noticeably stronger
  • junior IDs are becoming dependent on AI-generated structure
  • reviews are becoming lighter rather than sharper
  • the team talks more about prompts than about learning quality
  • AI is being used heavily for drafting but rarely for critique or audit
  • there is no shared method, only individual experimentation
  • leaders are celebrating adoption without examining instructional standards

These are all signs that the team may be using AI actively but not yet using it with sufficient maturity.

That distinction matters a great deal.

Because L&D leaders are not just accountable for efficiency. They are accountable for the quality of the team’s professional work.

The Leadership Mindset Shift

This is really the deeper point.

Leaders should stop asking only these questions:

  • Are we using AI enough?
  • How fast can we scale it?
  • How many people are adopting it?

Instead, they should start asking:

  • Is AI improving instructional quality?
  • Is it strengthening judgment or replacing it?
  • Are we building better review discipline?
  • Are weaker designers learning faster or just generating more?
  • Is our team becoming more consistent, thoughtful, and capable?

Those are better leadership questions.

Because they focus on what actually matters in an instructional design team: not tool excitement, but stronger professional performance.

The Larger Point

GenAI will reward organizations that know how to evaluate it intelligently.

If leaders use shallow measures, they will get shallow answers. They may think the team is progressing because usage is rising, while missing the fact that quality discipline is weakening underneath.

If leaders use stronger measures, they will see something more useful: whether AI is truly helping the team become faster, sharper, more consistent, and more capable.

That is the right standard.

Because the success of AI in instructional design should not be judged mainly by how often it is used. It should be judged by what it does to the quality of the work and the quality of the thinking behind the work.

That is what L&D leaders should be evaluating.

Instructional Design Meets AI – A Guide for Experienced IDs

LearnFlux 2026