My son sat for the Arizona state writing exam. The browser locked. Tabs closed. Phone gone. No notes, no search. Just a kid and a screen.
Then he walked out into the real world.
Work there doesn’t start with a blank page. A memo begins with research, maybe an AI prompt, then edits, fact-checks, and human review. A meeting becomes a transcript, a summary, an action list.
What counts as work is changing.
Assessment systems haven’t caught up. They assume a test must be secure to be valid. That’s wrong. A test can be perfectly locked down and still be useless. The issue isn’t cheating. It’s that tests behave like AI doesn’t exist.
The debate focuses on misconduct. Can a chatbot write an essay? Should schools ban AI?
Those are the wrong questions.
The real issue is the construct. Assessment specialists use that word to describe what we are actually measuring. Communication isn’t static. Constructs are maps of capability. If the terrain changes, you redraw the map.
How AI Changes What It Means to Be Competent
AI is not just a fancy spell-checker.
It changes authorship. It changes how we prove competence. Tool-free tasks show what a student can do alone. Stopping there confuses isolation with authenticity. You can preserve purity and lose relevance.
Xiaoming (Madeline) Xi, an assessment scholar, said it plainly.
“What we measure must adapt to learners’ AI-augmented realities.”
The goal is technology-mediated capacity. Knowing when to work solo. When to prompt. When to co-create, verify, or refuse.
Writing is one of the most common uses of AI, based on millions of Claude conversations.
Traditional tests look at the finished paragraph. Is the claim clear? Is the grammar right?
That tells you nothing anymore.
Polished prose hides the process. The meaningful evidence is upstream.
Did the student define the audience? Compare alternatives? Verify claims? Detect hallucinations? Reject bad output?
The writer is now an editor. A fact-checker.
Voice is what remains. It is the thing you preserve against the machine’s push toward smooth, agreeable sameness.
The Shift in Listening and Speaking
Listening has changed too.
Speech is recorded, transcribed, searched, converted to tasks.
Good listeners need to hear tone. Hesitation. Irony. What wasn’t said.
They need to spot when a transcript misses the point. When a summary flattens disagreement. When a follow-up turns ambiguity into false confidence.
Speaking?
That stays human.
Live delivery relies on timing. Credibility. Reading the room.
But the prep work? Research, rehearsal, translation, captions? All mediated.
Testing delivery alone misses the preparation. It misses the adaptation. It misses accountability.
AI hasn’t erased human skill. It has moved it.
Assessment must detect what fluency hides. Weak arguments. Fabricated sources. Distorted summaries. Plausible nonsense.
Schools need three types of evidence:
1. What students can do alone.
2. What they can do with AI.
3. Whether they can use it responsibly.
The last one is the hardest to fake.
Frameworks for Assessing AI Literacy
The 2026 AI LiterACY Framework from the European Commission and OECD gets this.
It treats AI literacy not as button-pushing. It’s the capacity to engage, create, manage, and shape AI. It weighs risks against benefits.
It moves beyond consumption to agency.
Assessment needs a ladder of readiness.
Foundational learners need protected space. They build vocabulary and reasoning without outsourcing the struggle.
Intermediate learners use AI for bounded critique. Counterclaims. Revisions.
Advanced assessments measure leveraging AI for complex problems. Verification. Audience adaptation. Rhetorical refinement.
The architecture should be dual-track.
Track one: AI-restricted tasks. Measure independent performance.
Track two: AI-inclusive tasks. Measure judgment under realistic conditions.
Imagine a student drafting alone. Then asking a model for a counterargument. Verifying its evidence. Explaining which revisions she accepted.
Or comparing a lecture to its automated transcript. Identifying what the record lost.
In these tasks, process is performance.
Look at the goals. The prompts. The drafts. The decisions to accept or reject.
College Board does this with AP Seminar and AP Research. They surface decisions. They trace the work. The aim? Access to thinking that the final product conceals.
Risks of Bad AI Assessment Design
There are three major risks.
First: cognitive surrender. Steven Shaw and Gideon Nave call it distinct from cognitive offloading. It happens when students adopt the machine’s answer without forming their own view.
The result is comprehension debt. More language produced than understood.
Second: tool fluency varies. Students have different “dialects” of prompting.
Third: homogenization. If you reward sterile polish, you train students to sound like robots.
These dangers argue for design, not denial.
Classrooms are laboratories. Low-stakes tests can experiment.
High-stakes systems controlling credits or admission should wait. They must move only when evidence shows they measure skill, not privilege.
What Remains Human?
What can’t a machine do?
It can draft the sentence. Summarize the meeting. Generate slides.
It cannot assume responsibility.
It cannot stand behind a claim. Answer for an omission. Empathize with an audience. Repair a broken relationship.
Judgment cannot be outsourced.
This is the accountable human remainder.
Ownership of purpose. Accuracy. Voice. Consequence.
Go back to my son in the locked browser.
The old exam asks: Can you do this alone?
Sometimes we need that answer.
But the world asks more.
Can you use a tool without being used? Can you preserve accuracy and voice when fluency is instant? Can you remain accountable for words partly shaped by a code?
The future isn’t nostalgic prohibition. It isn’t technological surrender.
AI can widen the construct of communication. Or it can hollow it out.
Design decides.
A school that refuses the new questions protects an illusion.
It measures the past.




















