This post motivates a technique I’m tentatively calling “Nyström grading”. It uses a kind of low-rank matrix approximation method with the intent to help me produce higher-quality and more thoughtful feedback than I could with traditional grading methods. You don’t need to understand physics or matrix theory to understand it. Anyone who has taken a college exam can likely relate to the steps involved. This is the BayLeaf blog, so it is of course going to involve AI, but perhaps not in the way you’d expect.
Imagine we’re in a lower-division physics course like UCSC’s “PHYS 5C Introduction to Physics III.” The catalog text reads: “Introduction to electricity and magnetism. Electromagnetic radiation, Maxwell’s equations.”
I don’t teach physics, and it has been perhaps two decades since I took an E&M course. I’m using lower-division physics as an example here because we all likely have more intuition about batteries and fingers than we do about the stuff my courses usually cover.
Imagine I’m teaching this course. Maybe I like to make my exams fresh each year so that they respond to the unique spin I’ve put on the field this year. This means my questions are likely to be a little unpolished, occasionally mistaken, and not yet tested on their target audience: the undergraduates. I might have one or two teaching assistants who can help proofread the exam, but these are exceptional (which is to say atypical) physicists who don’t represent my audience well.
We’ve done our best to design our exam: 10 questions where you need to show your work, given to 100 students in the course. Now we need to grade it.
Forget any attempt to use AI in grading right now, and just think about traditional grading with human eyeballs. Should we spread the grading between the TAs so that each grades 50 exams? Should we unstaple the exam pages so that one TA grades the first 5 questions and the other the last 5 questions? Should we do some more sophisticated multi-pass process with discussion among the teaching team? How do we give students feedback on their exams? We think the students, these days, mostly want to know their score. We, as physics educators, think they should care about our detailed, non-score feedback as well, assuming we could generate it. But how are we going to write 100 × 10 = 1,000 little blocks of text with any sort of consistency? How are we going to get feedback to them in time for it to feel relevant?
If we can figure out a clean solution for this, maybe we could apply the process to weekly quizzes or homework problem sets, not just one big high-stakes exam.
Now imagine the course is tiny and the exam is smaller too, but I’ve lost my teaching assistants. Maybe I have just 5 students and 5 questions. I’d probably approach grading by giving all of the exams a holistic skim before writing any scores or feedback. I might notice that one of the questions was ambiguously worded, that there’s a certain misconception that makes a string of wrong answers feel right in context, or that the per-student, per-question messages I’m going to give students are likely to involve some repeated blocks of text I’d like to copy-paste.
I might make some personal notes, some by question, some by student, that will help me when I eventually get around to doing the per-exam official grading and feedback writeups.
It sounds so cozy and wholesome, and so unrealistic in the current situation for lower-division physics (or any other discipline here).
I brought up electricity and magnetism for a reason. In an E&M course, you learn that we have this idea of “conventional current”. We draw our pictures with directional arrows that suggest current flows from the positive terminal of a battery (or other voltage source) to the negative terminal. Even though this has long been known to be the opposite of what’s happening with the physical electrons in a circuit, the convention is so strong that we speak quickly and informally: “the current goes like this”, gesturing with our hand.
Another relevant E&M convention is Ampère’s right-hand grip rule. Once you know which way the conventional current is flowing in a wire, you imagine gripping the wire with your right hand, pointing your thumb along that direction and curling your fingers around the wire. The direction your fingers curl shows you the (again, conventional) direction of the magnetic field lines. These are just conventions, but we teach them alongside other ideas in physics with such consistency that students can come to believe that they are as physical as the rest of physics.
With these two related kinds of discipline-specific conventions taken for granted, it is easy to write exam questions that look great to teachers and their assistant proofreaders who are intimately familiar with them. However, if a student misapplies a convention (which might or might not be what the exam question was designed to assess, often not), whether their answer is correct can flip, or even double-flip. They can be right for the wrong reasons, or wrong with right reasoning if you excuse the unconventional approach.
When we are grading these imperfectly designed exams (or accept that perfectly designed exams are something we shouldn’t reach for when other impacts are considered), we ought to be aware of how the imperfections interact with the student population. Often, those interactions are not observable until after the exam materials have been returned by students!
I’d love there to be some way for every student’s answer, in its full context of other answers to the same question, to potentially influence our interpretation of every other student’s answers to other questions. But it can’t come down to me, the teacher, reading every exam multiple times.
Now let’s bring it back to AI stuff. Lower-division physics (or lower-division anything else we teach at volume) tends to be well covered by the knowledge cooked into current LLMs. I’m interested in using LLMs, but I want to immediately knock away two misuses of them. First, let’s not chop up the exams into little pieces and feed them to an LLM for autograding one question at a time, separating a given response from other responses from that student or from other responses to that question by other students. This’d be wasteful, unreliable, blah, many problems. Second, let’s not attempt to just feed the entire grid of questions and answers for all students for all questions into the LLM and ask it to write the final evaluation grid. We’d probably get interesting cross-student and cross-question effects, but I’d suspect we’d have worse reliability as limited context precision gets students and their answers mixed up.
For reference, I’ve tried versions of these techniques on the materials for my own courses, and the experience of spot-checking results always convinced me that the two methods above are structurally unsound.
So what ought we to do instead? Take a sample of questions and a sample of students, and form a block of their responses. This will show some cross-question effects and some cross-student effects. Don’t ask the LLM for scores or student feedback; ask for notes to be accumulated at the student and question levels. Do this a few times with different samples, and then merge the notes. You’ll build up notes on student-question interactions not by populating a huge matrix, but through row and column summaries. Even with big courses (large n) and big exams (large q), we’re avoiding getting lost in any n × q product (and via subsampling we’re controlling which sparsified view of the n × q matrix is seen by the LLM, different every time).
In the physics example, a question-level note might say: “This question leaves the current-direction convention implicit. Check whether reversed answers reflect that ambiguity. Per-student feedback could comment on our ambiguous question design.” A student-level note might say: “Across the sampled responses, this student consistently treats current as electron flow. Their other reasoning appears consistent once that choice is accounted for. Penalize the mistake only in the first occurrence, then assume inverted convention.” These notes represent my plan for being more precise and consistent in my later feedback on a per-student-per-question basis. I can make the notes say whatever I want, but the point is that they are informed by what actually happened during the exam and how that compares with my pedagogical intent. They aren’t a context-free judgement of only the literal question or local response texts.
Now, when it comes time to give students specific scores on specific questions or to write up the per-response feedback text, you might or might not use AI, your teaching assistants, or your own eyeballs and wisdom. For today, I’m interested in the use of AI to discover, document, and refine some factored representation of that giant n × q matrix.
The name “Nyström grading” isn’t super well thought out. There’s a vague connection to Nyström methods for low-rank matrix approximation here, but the connection to loopy belief propagation and item response theory is maybe equally tenuous. The point is to get us thinking about ways AI can be involved in grading that don’t boil down to a simple “human vs. machine” dichotomy. In the interest of broad accessibility, I’m interested in scalable methods, but at a fixed scale I’m interested in ways of improving the quality and thoughtfulness of my feedback for students under a finite per-student budget of labor hours, shrinking under austerity, available to author that feedback.
This blog post is timely for me because I ran a version of Nyström grading for a quiz in my software engineering course this past week. I didn’t give students a clear enough basis to judge the boundaries of CI/CD, some collapsed the idea getting their game to enter Unity’s Play Mode and building a redistributable binary, and so on. My method was a little different and much messier than sketched in this post, and it took almost as many hours to execute as I’d have done eyeballing the responses directly in the traditional mode. So, this blog post is mostly my attempt to clean up the ideas and make them transmissible so others can try out their own, discipline-specific variants.
Perhaps you are wondering: Did he feed everyone’s responses to a big commercial AI provider, along with his rubrics and stuff, just to feed the beast that will sell all of that work back to us in the form of expensive edtech software next year? Nope. We have nice tools here at UC Santa Cruz that you can’t buy, at least in the constellation I used. The insights gleaned from the analysis are stored primarily on my laptop and shared directly with my students and teaching team on the course Canvas, and with you here via the BayLeaf blog, not snarfable by unintended third parties.
Ah, I didn’t give you any coverage of related work here because this was just a quick writeup. Please don’t let that stop you from spidering out from there. Tell your agent: “Help me to some related work search relative to a recent blog post: https://blog.bayleaf.dev/p/nystrom-grading I want to know about other people doing something similar either in the current LLM moment or previously. Emphasis on peer-reviewed academic work.” The work is out there, and you’ll probably learn more by reading that work specifically in relation to this blog post than if you encountered in isolation. Check it out. Contextual diagnosis and machine-assisted grading are ancient topics, what matters is whether you have the tools and experience to set you up to put these into practice in your own teaching next week.

