Skip to content
HN On Hacker News ↗

How I changed teaching after AI managed to do all my homework assignments

▲ 305 points • 288 comments • by azhenley • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

0 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,675
PEAK AI % 0% · §1
Analyzed
Sep 26
backend: pangram/v3.3
Segments scanned
1 windows
avg 1675 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,675 words · 1 segments analyzed

Human AI-generated
§1 Human · 0%

Around 2021, well before ChatGPT launched, Vincent Hellendoorn suggested I try GPT-3 on the reading quizzes in my course. It produced convincing answers passing our rubric without actually seeing the assigned paper. At the time, I changed nothing. Five years later, AI agents could do all my assignments and I have redesigned most assessments in that course, even though what I want students to learn has barely changed. The strategy is always the same: No longer test understanding with anything that is done at home and instead focus on interactions with a TA, on exam, and on a video demo. Some of these changes violate evidence-based best pedagogy practices, and I made them anyway.For the last couple of years, I have mostly taught the course Machine Learning in Production, an upper-level course on building production-ready software around ML models with a heavy focus on MLOps, usually with 100 to 170 students. These days one common question when talking to other educators is how we have changed teaching in the age of generative AI and coding agents, so let me outline what we did.We have shifted covered topics with changing AI innovations and tools, but I barely touched the overall learning goals. I am fortunate that this is not an intro course and that the learning goals are not about writing code or using specific tools; they are about engineering tradeoffs, anticipating and mitigating risks, and teamwork. I think these are skills still worth acquiring, even if some can be simulated and offloaded to a model. (Revising an intro course or a traditional software engineering course likely would shift learning goals much more.)Also possibly important: We give students permission to use AI in all settings, in any form, without attribution, except for written and oral exams. We even encourage the use of AI tools in many places. I do not think policing AI is feasible even if we wanted, and more importantly I do think that students need to learn responsible use of these technologies anyway.Let’s start with this point upfront, since it is more important than what we actually do: Unfortunately, AI is actively undermining several evidence-based teaching practices (e.g., see How Learning Works and The ABCs of How We Learn). For example, the evidence favors frequent low-stakes assessments with feedback (e.g., homework, quizzes) over few high-stakes ones (e.g., exams) – but AI is undermining practice in low-stakes settings and pushing us more toward exams.Similarly, I always provided a safety net where students can make mistakes and resubmit a limited number of assignments to regain lost points (a core recommendation of specifications grading and grading for equity to focus on learning outcomes, not the process), but we felt that this process was abused with AI: first submit a generated assignment solution without thinking and only look at the issues raised in grading for a resubmission (the typical story of externalizing the cost of AI use). In response, we have since taxed resubmissions with a 10% penalty.Also in-class interactions allow engaging with materials in an early low-stakes setting, but with AI I have seen many student groups offload the discussion questions to a model. Pen-and-paper submissions could fix this, but aside from a higher grading workload, it would also raise stress for students, take away from the low-stakes environment, and delay feedback.In general, this is a balancing act and I tend to err on the side of keeping low-stakes repeated interactions even though it can be abused. Yes, some students will get through the class without much deep learning, but it provides a better environment for those students who want to learn. I don’t want to get back to the model I’ve experienced during my own studies in Germany with mostly optional homework and a single exam at the end of the semester that was responsible for 100% of the grade in the course. This was nice for students who were self-motivated and good at learning for an exam (like me, I guess), but had failure and drop-out rates of 50 to 80%.Now for actual changes in the course: I have given up all parts of assignments that required a written text answer. I still ask for reports that describe a solution and link to the relevant code fragments, but that’s just for navigating their solution and I’m fine with receiving AI generated documents for that. In contrast, reflection documents, like “What were challenging parts?”, “How would you improve teamwork?” or in reading quizzes “For scenario X, identify one plausible data quality problem you might expect that relates to one of the four data cascades discussed in the paper…” have become pointless and can be entirely delegated. Short of hiding the evaluation rubric, I can see no way of stating what I expect in a good answer that cannot be completely offloaded to an LLM.For written reflections, which I used to have as a part of pretty much every assignment, I now shifted to in-person interactions with a TA. After every assignment, each student needs to schedule a 15-minute meeting with a TA to answer a couple of questions in a live conversation (apparently Stanford cs221 is evaluating the same kind of approach in a controlled experiment this semester). I still share the reflection prompts in the assignment as examples of the kind of questions we ask. Students can still generate an initial answer with an LLM, but they may need to memorize parts of it, and we try to challenge them with follow up questions. The check-in meetings are part of the assignment and currently worth 20% of the assignment points, graded pass/fail. Students can try again if they fail and I encourage my TAs to have fairly high standards – we usually fail quite a few students on their first attempt.There are drawbacks to this design, but overall I am happy with the tradeoffs: Penalty-free retries reduce fairness concerns about TA grading of oral interactions; a more strict TA costs a student time, but not points. Oral check-ins demand more from students with anxiety, but so do written exams, and formal disability accommodations can provide a path in both cases. In fact, professional communication about technical work is a learning goal and oral check-ins train this more than written reflections. Regarding scale: We run the course at a 20:1 student-TA ratio with about 10h of work per TAs per week (fortunate, I know), so the check-ins amount to roughly 300 minutes per TA every two weeks, which is workable.For reading quizzes, I just gave up. I did not think doing in-person pen-and-paper quizzes in class would be worth the stress and the needless memorization work that those would be causing. I actually kept online reading quizzes around for a long time just to signal that I wanted students to look at the paper, fully understanding that most would just ask an LLM. These days, I still assign readings, but only half as many and without any points attached. Instead, I try to integrate lessons from the readings into in-class discussions. Still most students do not do the readings and just ask an LLM when we get to that point in the class (so nothing changed on that front), but those that do might get more out of it.Minor note: We observed that some students used AI during live discussions over Zoom (e.g. Cluely) and we will likely only offer in-person checkins in the future in response.We have weekly labs that are low-stakes small tasks to explore new tools (e.g., Kafka, Grafana, Docker, Weights and Biases). These tasks are necessarily scoped small and need to provide some help to students starting out – so they are obviously easily automateable by coding agents. We again rely on in-person check-ins – show the TA evidence that you completed the task and be able to answer a few questions, graded pass/fail. Again, we frequently send students back to read more documentation (or let their chatbot summarize the relevant part) and let them try again without a penalty until the time of the lab session runs out.For assignments, we use the same check-in with a TA discussed for reflections to let them explain part of their technical solution. In teamwork, we have longer debriefing sessions after each milestone (30-60 min per team). We award “beyond-the-comfort-zone” bonus points to a team if the TA can ask any team member to explain any part of the implementation. (Yes, I know, bonus points are a scam. I use them anyway. Sue me.)Also having students produce a short video demoing a feature they implemented worked really well for an assignment to extend a web application, because it required the feature to actually work with a user interface in a real workflow. I think producing videos for other parts could also work, as long as it’s not just reading an AI-produced script, but actually grounded in some technical work.As I see in many other courses, we also shift more points from activities done at home (e.g., homework) to activities done in the classroom (e.g., exams, participation). Exams are now worth 25% instead of 15%, and I suspect I will raise this further in the future. The debriefing is 10 to 20% of the homework and group work grade. Still the majority of points are associated with homework and group work done at home and most students get full or nearly full credit, but the main grade differentiation now comes from exam grades.I have not yet introduced graded in-class pen-and-paper quizzes that many other instructors now use, but it is an option. I prefer debriefing with a TA for now.As students can more easily produce large amounts of code and lengthy text documents, traditional manual grading has become more tedious (the typical asymmetry of lower production costs without lowering manual review costs). The turning point for me was a year ago, when a TA shared how he felt silly grading a solution where the commit message included “Authored by Claude Code.”We have since built infrastructure to autograde code and written reports of homework with LLMs