AI for teachers

The AI Mirage at University: Lots of "Artifacts" and Little Innovation

Cover for The AI mirage at university: lots of artifacts, little innovation, with the UABC study figures: artifact production 76.7%, critical thinking, verification and ethics 11.3%, iteration 3.5% and personalization 1.0%

I have spent more than twenty years in classrooms, and almost every week I hear the same promise: generative artificial intelligence is going to "revolutionize" the university. That is why I stopped at a recent paper by Karla Karina Ruiz Mendoza and Eilen Oviedo González (2026) that reaches an uncomfortable conclusion: much of what gets called innovation with AI is producing the same things, in more formats and faster. My own summary is blunter: generative AI (GenAI) is being used as a glorified typewriter. But before assigning blame I did what I ask my students to do: I verified. I got the full paper, which is open access, went through its tables and compared what it says with what is usually said about it. The result is more useful, and more cautious, than the headline.

In this article I tell you what the study says and what it does not, what the wider evidence shows about AI in higher education in Colombia, Latin America and the world, and I leave you a concrete proposal (labeled as a proposal, not as proven truth) for assessing the process and not only the product, with a rubric, a logbook template and an oral-defense protocol you can use tomorrow.

What the study really says about AI in higher education

It is titled "From artifact production to pedagogical innovation: teaching strategies with GenAI in higher education" and appeared in issue 44 (special issue, September 2026, pp. 40 to 49) of the Revista Iberoamericana de Tecnología en Educación y Educación en Tecnología, from the Universidad Nacional de La Plata (Argentina). The authors work at the Universidad Autónoma de Baja California (UABC), in Mexico. It is a qualitative, exploratory and interpretive study, using content analysis and descriptive percentages, with no claim to generalize.

The corpus is 600 teaching strategies with a valid description, designed by 186 UABC instructors and recorded during an institutional teacher-training course on GenAI. Instructors came from all areas: 22% engineering and technology, 22% economics and business, 16% arts and humanities, and the rest from social sciences, health, natural and exact sciences, law and education.

Chart with the figures from the study of 600 strategies by 186 UABC instructors: artifact production 76.7%, guided practice 15.5%, critical thinking, verification and ethics 11.3%, feedback 10.7%, planning 7.8%, automation 5.8%; innovation indicators such as collaboration 10.3%, verification 9.3%, iteration 3.5%, ethics 2.8% and personalization 1.0%; and assessment: product or project 49.0%, rubric 33.5%, no criteria 22.5%, reflection 5.3% and exam 3.3%
What instructors declared: artifacts abound, while verification, iteration, ethics and personalization are scarce. Source: Ruiz Mendoza and Oviedo González (2026), tables 2 to 4.
Evidence: the study's main figures (N = 600 strategies)
What was measuredResult
Main toolChatGPT 29.2% and Landbot (no-code chatbots) 17.8%; then Picsart 11.8%, InVideo 11.7% and SlidesAI 6.7%; "others", 13.5%
More than one tool in the same strategy22.2%
Instructional integration: artifact production76.7% (460 of 600): text, slides, image, video or chatbot
Critical thinking, verification or ethics11.3%
Explicit innovation indicatorsCollaboration 10.3%; verification and attribution 9.3%; prompting as a skill 4.8%; iteration and improvement 3.5%; ethics 2.8%; personalization 1.0%
Declared assessmentProduct or project as evidence 49.0%; rubric or explicit criteria 33.5%; no instrument or workable criteria 22.5%; reflection 5.3%; exam or quiz 3.3%

A strategy can fall in several categories, so the percentages do not add up to 100. Source: tables 1 to 4 of the paper.

What the study does support

Three things, with numbers: activities that produce artifacts dominate; practices such as verifying, iterating or discussing ethics appear rarely; and assessment revolves around the product. The authors sum it up themselves: innovation "is expressed more as diversification of formats than as systematic redesign of assessment, verification and personalization processes". My thesis, which I stand by, is that a flawless essay, report or piece of code no longer tells you, on its own, who did the thinking.

Where it needs nuance

I say this with respect for the text that started this reflection, and for my own: some formulations stretch further than the data.

  • "The universities" are, in fact, one Mexican public university. The authors warn that the study "is limited to one institution and a specific moment of adoption", so it should not be generalized without considering discipline and local policy.
  • "What professors do" is, more exactly, what they wrote in a course form. It represents pedagogical intent, not classroom implementation or learning effects, and it does not include what students did.
  • "Absolute dominance of ChatGPT and Landbot": together they add up to 47%, not everything. The authors also acknowledge that several tools were presented in the course and instructors picked the most familiar ones, so part of the concentration is an effect of the course itself.
  • "Critical thinking is absent": the study says "less frequent", and it only counts explicit mentions. A professor may demand source checking in class and not write it on the form. The authors speak of "a gap in pedagogical explicitness rather than an absence of normative concern".
  • "Conventional rubrics": the paper does not call them that. In fact, it argues for rubrics that include attribution, verification and an explanation of AI use. The harder number is a different one: almost one in four strategies leaves no operational assessment criteria at all.
  • "Assessing the final product is absurd" is my thesis, not a conclusion of the study. The authors say something more nuanced: the validity of traditional tasks is under strain and the process should be made visible through prompt records, logbooks, source checking and successive versions, which is exactly what I propose below.

A note on sources: next to my original draft circulated a phrase attributed to "Revista Mundo Empresarial (2026)" that I could not find in any publication, so I do not cite it. I prefer one fewer quote and one more verification.

Why assessing only the product stopped working

The UABC study is a portrait of intentions. To understand the underlying problem we have to look at what students do and what we know about learning.

  • Use is almost universal. In the UK, the 2026 HEPI and Kortext survey (1,054 undergraduates) found that 95% use AI in some way, 94% for assessed work, and 12% pasted AI-generated text as is into assessed work. In 2024 the figures were 66%, 53% and 3%. And 65% say assessment changed significantly because of AI, while some students express anxiety about being falsely accused.
  • Practicing with unguarded AI can be costly. In an experiment with almost 1,000 high school students in Türkiye, published in PNAS (2025), those who practiced with an unrestricted chatbot improved their practice performance by 48% (127% with the tutor version) but scored 17% worse than the no-AI group when access was removed for the exam. The "tutor" version, which gave hints instead of answers, eliminated the harm, though it produced no improvement. These are high school math students, not university students: a signal, not a verdict.
  • Performing is not learning. The OECD's Digital Education Outlook 2026 concludes that, without pedagogical support, delegating tasks to AI improves immediate performance but not real learning.
  • With caution: the MIT preprint "Your Brain on ChatGPT" (2025) had 54 participants, only 18 of whom completed the fourth session; it has not been peer reviewed and is limited to essay writing. It helps formulate questions, not write headlines.
  • Detectors are not the way out. Liang and colleagues (2023) showed that seven detectors flagged, on average, 61% of 91 TOEFL essays written by non-native English speakers as "AI-generated". And OpenAI withdrew its own classifier in July 2023 because of low accuracy: in its initial evaluation it caught only 26% of AI-written text and falsely flagged 9% of human text.
Bars from the HEPI survey of UK students in 2024, 2025 and 2026: use of AI in some way 66, 92 and 95 percent; use in assessed work 53, 89 and 94 percent; AI text pasted as is 3, 8 and 12 percent; and four cards: 86 percent of 3,839 students in 16 countries use AI, 84 percent of Colombian students use it frequently and only 35 percent go beyond basic skills, 83 percent of 1,681 faculty worry students cannot critically evaluate AI output, and minus 17 percent on the AI-free exam for students who practiced with an unguarded chatbot
Almost all students use AI. The cards come from different surveys and studies and are not comparable. In 2026 HEPI adjusted the 2025 figure to 89% (it was 88%).

My conclusion is not "let's monitor more" but "let's design better". If neither bans nor detectors solve the problem, what is left is redesigning what we assess and how. It was the same with phones in the classroom: banning is not enough, we have to teach.

Colombia, Latin America and the world: three snapshots of the same film

Colombia: heavy use, little training and guidance without binding force

According to a GAD3 study for Planeta Formación y Universidades (presented in Bogotá in November 2024), 84% of Colombian higher-education students use generative AI tools frequently, but only 35% have the skills to go beyond a basic level. At the official level, the Ministry of Education's Decalogue of AI for higher education (August 2026) respects university autonomy and asks, in its sixth point, for assessment processes "that make it possible to observe thinking processes and not only final products". It builds on CONPES 4144 of 2025. It is guidance, not a binding rule; and as for a specific AI law, I only found bills in progress (Bill 043 of 2025 in the Senate), with no confirmation that they have been passed. At the institutional level, the Universidad de los Andes published guidelines in October 2024 that ask students to declare how AI was used and let professors set a scale of use per activity. And there is a national "secure lane": the Saber Pro test was taken in April 2026 in person, on paper, although it measures generic skills and does not replace assessment in each course.

Latin America: reactive research and guidelines that are only arriving

A systematic review of 35 studies (2023 to April 2025) on teaching with GenAI in Latin American higher education found that the dominant themes are teacher training and academic integrity (54% each), while technological inequality (17%) and the lack of institutional policy (20%) stay in the background. The UABC study fits that picture. UNAM, in Mexico, published in 2026 guides on generative AI in assessment, with a sentence from Melchor Sánchez Mendiola, head of its evaluation office, that I share: AI "does not destroy assessment, it forces us to improve it". And the region already has instruments such as CriticalAI, which we will see below. Since 2023, the IESALC-UNESCO guide on ChatGPT in higher education has served as a regional starting point.

The rest of the world: reform assessment, don't just police it

In 2024 the Digital Education Council surveyed 3,839 students in 16 countries: 86% use AI in their studies and 58% feel they lack sufficient knowledge. In EDUCAUSE (2025), only 39% of respondents said their institution had AI acceptable-use policies (a year earlier it was 23%). Australia took a concrete step: the regulator TEQSA proposed in 2023 that assessment prepare students to take part ethically in a society with AI and that learning be judged through multiple approaches; the University of Sydney turned it into its "two-lane" model. UNESCO published its guidance on generative AI in education in 2023 and, in 2024, the AI competency frameworks for students and teachers. And oral defenses are returning to classrooms: an AP report describes their comeback in US universities, with cost as the main obstacle.

Four viewpoints: professors, university leaders, students and families

Professors

In the Digital Education Council's 2025 global survey (1,681 faculty in 28 countries), 61% had used AI in teaching, but 83% worried that students cannot critically evaluate what AI produces and 80% said their institution is unclear about how to apply it. And 54% thought assessment methods require significant changes. I did not find a representative survey of Colombian professors, and I would rather say so than invent a percentage.

University leaders

Theirs is a dilemma of scale and risk. An oral defense or a layered task costs teaching time, and teaching time is money. But a degree nobody can say what it certifies costs more. The Ministry's Decalogue and the Uniandes guidelines suggest the path: clear rules per institution and autonomy per course.

Students

According to HEPI (2026), only 36% feel their institution encourages them to use AI and 48% think their teaching staff help them develop AI skills, although 68% consider those skills essential for their careers. Among the comments in the report there is a fear we should take seriously: being accused of using AI simply for writing well.

Families

I did not find a Colombian survey of what families think, so what follows is my reading and not data: someone who pays tuition wants a degree that means something to an employer, and does not ask the university to ban AI but to teach how to use it with judgment. For them, process assessment is a guarantee that their child learned, not just that they handed something in.

A proposal for assessing the process: four pieces (not a proven recipe)

What follows are design proposals. They rest on the literature and on teaching craft, but I know of no experiments proving this set. Try them in one course, measure what happens and adjust.

Diagram: assessing only the product (prompt, AI writes it, document handed in, grade whose work) versus assessing the process in four layers (compare, audit, document, defend) and, below, two lanes: lane 1 secure, in person and supervised, and lane 2 open, with AI allowed and declared
A four-layer, two-lane design proposal. It is a proposal, not proven evidence.

1. Layered tasks: compare, audit and conclude

Instead of "write an essay on X", ask the student to use AI to generate three positions on X, and assess their analysis: how they compare them, which fallacies and biases they detect, which claims they verify and what they conclude using readings the model did not give them. That way AI stops writing for the student and becomes their counterpart.

Assignment sheet (example, Business Administration course). Topic: should a Colombian small business adopt permanent remote work?

Part 1 (with AI, 20 minutes). Ask the model for three positions (for, against and in between) with arguments and sources. Save the whole conversation.

Part 2 (without delegating, 800 words maximum). Compare the positions using criteria you define; name at least two fallacies or biases and quote the passage; verify five factual claims in primary or indexed sources and mark which turned out false, nonexistent or unverifiable; add two readings that did not come from AI; write your conclusion and the variables the model left out.

Part 3. Hand in your prompt log and your AI-use declaration (level 4 of the scale below).

Part 4. Defend your analysis in 8 minutes, without slides.

Proposed rubric (each criterion scored from 1 to 4)
Criterion and weight1 Beginning2 Developing3 Achieved4 Outstanding
Comparing positions (25%)Summarizes the three without comparingCompares with a list of pros and consCompares with explicit criteria and locates the real clashReveals hidden assumptions and reframes the problem
Fallacies and bias (20%)Detects noneDetects one without quoting the passageDetects two or more, names and quotes themAlso explains how they change the conclusion and how to fix them
Source verification (20%)Accepts data and references uncheckedVerifies some; confuses primary and secondary sourcesVerifies five claims and flags the false or nonexistent onesAlso adds two own readings that qualify the argument
Metacognitive log (15%)Missing or genericCopies prompts without analyzing themRecords prompts, AI errors and corrections with the reason whyShows how their way of asking changed
Oral defense (20%)Cannot explain their own claimsRepeats the text and fails the variationExplains, justifies and answers the variation with reasonsArgues under pressure, admits limits and cites what the AI left out

One design decision I suggest: let the defense cap the product's grade. If the student scores 1 on the defense, the document cannot exceed 2.5 out of 5. That way the defense is not decoration. If you prepare these tasks with AI, the AI Kit for Teachers includes recipes by subject for asking for contrasting positions, building question banks and drafting rubrics, always with your judgment having the last word.

2. Metacognition and self-correction: the prompt log

One source needs clarifying here. The journal RELIEVE published in 2026 the instrument CriticalAI, by Katherine Sandoval Perdomo and Luis Gibran Juárez Hernández: 48 items, six dimensions of Facione's model (interpretation, analysis, evaluation, inference, explanation and self-regulation), with items such as checking information obtained from AI against reliable sources or identifying its biases. But it is a self-report questionnaire validated in a pilot of 81 health-science students at a private university in Santiago, Chile, without factor analysis yet. It works as inspiration, not as a ready-to-grade rubric. For that, use the log:

Prompt log template (one row for each important exchange)
No.Prompt (verbatim)What the AI answeredWhat I verified and with which sourceWhat I fixed or discarded and whyWhat I would change in my next prompt
1"Give me three positions on permanent remote work in Colombian small businesses"Three positions with two figures and one citationI looked up the figures at the official source; one does not appear and the citation does not existDiscarded the figure and the citation; used the data I did findAsk for links and separate facts from opinions
2(illustrative example to complete)

3. Oral defense: an 8 to 10 minute protocol

  1. Opening (1 min). The student summarizes their thesis, with no slides or notes.
  2. Depth (3 min). Three questions about their analysis and sources.
  3. Variation (3 min). You change a figure or an assumption and ask them to rethink their conclusion.
  4. Process (2 min). "Show me a moment when the AI got it wrong and how you noticed."
  5. Close (1 min). The student self-assesses in one sentence and you score with the rubric on the spot.

A starter question bank: which part of this text could you not defend without the AI? Which claim was hardest to verify? If the opposing position were true, what evidence would show it? Which variable did the model leave out and why does it matter? If this figure changes, does your conclusion still hold? How would you explain it in 30 seconds to someone with no background? Which of the sources you cited did you read in full? What would you do differently in your first prompt?

With 40 students that is six or seven hours, so you can run sampled vivas (a randomly chosen part of the group, announcing beforehand that anyone can be called) or five-minute mini defenses. Plan reasonable accommodations for anyone with anxiety, a stutter or a hearing disability (more time, written questions, a support person). The PIAR under Decree 1421 of 2017 is for preschool, primary and secondary education, but its logic applies at any level; if you teach in a school, PIAR with AI helps you build one, and the first is free. And to schedule slots and keep attendance, an attendance sheet in Excel works well.

4. AI-use declaration and two lanes: what holds the system together

None of the above works if the student does not know what is allowed. The Universidad de los Andes adopted in its guidelines a five-level scale proposed by Perkins, Furze, Roe and MacVaugh (2024), which the professor can set for each activity:

AI-use scale per activity (adapted from the Uniandes guidelines, 2024)
LevelWhat is allowedWhat the student must declare
1No generative AINothing
2Explore ideas and structure; generated material does not go into the workThat they used it to explore
3Improve the clarity of the student's own ideasHow they used it, in a footnote or appendix
4Complete some elements with human evaluationGenerated material in quotation marks, its quality and relevance, and a prompt appendix
5Full use throughout the processNo need to distinguish own from generated

And a two-lane architecture, inspired by Sydney: lane 1 is secure (in person and supervised: in-class exam, oral defense, lab practice) and assures that learning outcomes were reached; lane 2 is open, with AI allowed and declared, and teaches thinking with the tool. For the secure lane you need exams that cannot be solved through a chat: the AI Exam Generator creates several versions of the same exam with an answer key, and you can try it for free in the demo. I explain it in this article.

What you can do this week: professors and students

If you are a professor

  • Run your most "essay-like" assignment through an AI and see whether your rubric would pass it; if so, you know what to redesign.
  • Write the allowed level of AI use (1 to 5) in the syllabus and in each assignment.
  • Add one layer to a single task: three positions and an audit of five claims.
  • Ask for the log using the template above and grade only what you can read in ten minutes.
  • Run a five-minute defense with a random sample of your group.
  • Reserve an in-person exam with different versions for the secure lane.
  • Never use a detector as the only proof, and never to accuse.

If you are a student

  • Always declare how you used AI, even when nobody asks.
  • Save your whole conversation: it is your evidence of work and your defense against an accusation.
  • Verify at least five facts from each answer against primary sources, and be wary of "perfect" citations.
  • Ask the AI "where could you be wrong?" and check the answer against a real source.
  • Ask it to quiz you instead of solving things for you: that is the difference between a tutor and a crutch.
  • Practice explaining your work out loud, without a screen, before handing it in.
  • Do not paste personal data, yours or anyone else's, into a chat.

If you want the other side of this coin, what AI promises and does not deliver, I collect my analyses in the AI for teachers section, together with "AI won't replace you. Whoever masters it will". And if you wonder why the format matters less than the quality of assessment, see online versus in-person education. The same goes for formulas: if a generation asks for formulas without learning to read them, it signs off numbers it does not understand.

Tools to assess the process, not just the product. The AI Kit for Teachers includes tested recipes by subject, aligned with the Colombian curriculum, for 60,000 Colombian pesos per subject. The AI Exam Generator creates different versions of the same exam for your secure lane, with plans from 29,900 pesos. PIAR with AI lets you build one PIAR for free and then choose packages of 5, 10 or 20 plans. None replaces your judgment: they are made so you can spend it on what matters.

See the AI Kit for Teachers

Frequently asked questions

Does this study prove that universities are not innovating with AI?

No. It analyzes 600 strategies from 186 instructors at a single Mexican university, recorded in a training course, and it reflects intentions, not classroom practice or learning results. It is a valuable signal and consistent with other reviews, but not proof about "the universities".

Should AI be banned in university assignments?

Not necessarily. Banning without being able to verify only punishes those who comply. A more realistic way out is to combine a secure lane (in person, no AI or controlled AI) with an open lane where AI is allowed, declared and the process is assessed.

Do AI detectors work to tell whether someone cheated?

Not as the only proof. In a 2023 study, seven detectors flagged on average 61% of some essays written by people whose first language is not English as AI-generated, and OpenAI withdrew its own classifier because of low accuracy. A conversation with the student, their log and an oral defense are fairer.

Is an oral defense feasible with groups of 40 students?

Yes, with adjustments: five-minute vivas, defenses for a random sample or defenses in pairs. It is not free in time, but it is the most direct way to check that the student understands what they handed in.

Isn't assessing the process much more work?

At first, yes: you have to design tasks, rubrics and templates. Afterwards they are reused, and AI can help with feedback drafts, as long as the decision about what counts as evidence of learning stays yours.

Food for thought

If an AI writes in 30 seconds an essay that passes your rubric, is the problem the student who used it, the professor who designed the task, or a university that for decades called "learning" what a machine now does? And if the oral defense is the answer, are we willing to pay its cost, with fewer students per professor and more time to listen to them, or will we keep grading documents nobody knows who wrote?

Was this article helpful?

React and share it with someone who might find it useful.

Press and hold to choose another reaction

Be the first to react

Prefer to have it ready?

More guides in AI and technology for teachers

Use ↑ ↓ to move, Enter to open and Esc to close.