University of Wisconsin–Madison

AI Showcase: Testing AI with Peer Evaluation

David McHugh

Instructor

David McHugh, M.A., Teaching Faculty and Instruction Manager, Information School

Course

Library and Information Studies 408 Generative AI: Strategic Application, Evaluation, and Critique

Assignment

Assignment instructions and rubric

Students are asked to design and implement a multi-trial test of an AI tool on a task relevant to their career. They write a results brief and reflect on the test’s limitations. They then run and evaluate tests designed by two classmates.

A theme of my generative AI course is: “Wonder if AI can do this? Test it.” We do this in a variety of ways and formats, but I wanted to devote some real time and value to it, so I’ve been growing this two-part assignment.

The primary learning goal is that students should practice the thinking and process for evaluating automation tools in different use cases. Specifically, this means strategically deciding what successes and failure results look like and how they can be assessed. A secondary learning goal is evaluating others’ work with constructive feedback.

This assignment builds on in-class exercises of assessing various tools and use cases involving generative AI (a particularly messy area to assess). I prefer assignments to be shared with the rest of the class, and this one’s a particularly good fit for peer feedback.    

Results

The first semester I used it, students often chose trivially basic tasks to test (e.g., AI for writing an email). Those were obvious (which still can have value).

Their work dramatically improved when I added an emphasis on novel tasks.  Students were creative and thoughtful, testing topics as varied as local laws about keeping chickens and translation in specific dialects of Spanish.

I’ve begun updating it for fall 2026; it’s still a work in progress (as are all my assignments in our AI era).   

Contact

dtmchugh@wisc.edu
LinkedIn