Pull-Based Inverse Classroomwork in progress

Concepts arrive before class. Difficulty arrives before class, from the students. The class hour is made from the difficulty.

Jake Benhart with Dr. Michael G. Kay · NC State University · College of Engineering · Operations Research

This is a first cut, and we are still building it. One course, one term, six students. The page is here so you can see the whole shape of it and tell us whether it could work in your subject — that conversation is the reason it exists.

The idea

An inverted classroom moves the lecture out of the room. This moves something else out too: the question of what the hour is for.

Students meet the concepts on their own, before class. Then, still before class, they work an example and record what they could not resolve. That record is the input to the class hour. The hour is not a prepared lecture delivered again. It is built, that week, out of what this particular group actually got stuck on — less time on what they already understand, more on the breakpoints they hit.

The word pull is the point. Nothing is pushed at a student on a schedule fixed in advance. The material a student receives is the material their own questions earned, and the difficulty the class works through is difficulty the class supplied.

The barrier to teaching this way was never desire. It was effort. Running a class from what students bring means reading what they brought, every week, before every meeting — and then building something from it. That is the part that now has machinery behind it.

Where the grade is earned

This is the part most people want answered first, so it comes first. Take-home work is done with AI and carries completion credit. Most of the grade is earned in class, on paper, by hand.

WeightComponent
60%Three examinations, in class, closed computer
25%Six projects — about 5% for the submitted work, about 20% for the in-class assessment of it
15%About eight homeworks — about 5% for the submitted work, about 10% for the in-class assessment

An AI tool can produce a take-home solution, so a take-home solution cannot certify what a student can do. An assessment written by hand, in the room, can.

That single move resolves most of what makes AI difficult for a course. Students are free to use it where it helps them learn, because the thing being graded is what they can do without it.

How a week runs

Before the lecture — the staff, once per lecture

  1. Draft the brief from the lecture's own text and nowhere else: which examples to rehearse, which validation check each one exercises, and the misconceptions to expect.
  2. Run a validation pass that reads the material the way a student will receive it. A clean self-test is not this pass and does not substitute for it.
  3. Publish through a tool that refuses to publish anything unvalidated. That ordering was bought the hard way: two briefs were published and then validated, and the pass could only decide how fast to repair a leak already in front of students.

Before the lecture — the student

  1. Read the lecture, watch the recorded overview, run the companion scripts.
  2. Run the review with an AI coding agent: one worked example at a time, answering before being told.
  3. Record what they could not resolve, and which validation checks they reached for.
  4. Push it. The repository is the submission. There is no upload step.

The hour itself

A pre-class report says what the class answered well and which sections to revisit, with every answer checked against the brief and mapped to a lecture section. The hour opens with what nobody resolved. No AI in the room.

What outlasts it

Each student has one profile file, rebuilt from their whole record every time rather than appended to, so it costs the same to read in week 15 as in week 2. It tallies which validation checks they have reached for and which they never have, by name, so a gap has a name. It suggests an order to study in. It gates nothing, and everything the student writes in their own notes survives every rebuild.

GitHub as the course engine

Everything moves through Git, which is doing more work here than it first appears. Every change is dated and attributed. Any earlier state of any document, student or instructor, still exists. A commit history shows both how work was done and when.

RepositoryDirectionHolds
materialsread-only inlectures and companion scripts
handoutsread-only inskills, briefs and sheets, once published
workread-write outone private repository per student
feedbackread-write outpreviously graded content

Students pull new material and push submissions. The commands they run update themselves from the same place, so nobody is ever told to reinstall anything.

Building the material

The raw material for a flipped course is often already sitting there. Ten years of class recordings is a corpus, and a synthesis across four recordings of the same topic produces something no single recording contains.

The workflow that turns a recording into a lecture keeps the instructor's own voice visible: the raw transcript passage stays on the page in grey beside the AI's gloss of it, each carrying a reference number, and the instructor accepts or rejects by number. Roughly two thirds clear on the first pass. The binding constraint is the speed of that loop, not its quality.

Once material is out there you are not done with it. When the worked example titles changed, about forty title changes were applied across every lecture and republished in half an hour.

One standard, taught and enforced

Two things the lectures teach are also what the machinery runs on, which is what keeps the tools from becoming a parallel curriculum.

A six-slot model format. Every model is stated the same way before any symbols or code: what is minimised or maximised, what is solved for, what it is subject to, what it returns, and what is assumed. A constraint restricts the solution; an assumption restricts the world. The model is a file, not a conversation, so it cannot erode as a chat window fills. A script checks the form — and says in its own output that form is all it checks.

Nine validation checks. Three screens on every result, one expectation on anything used in a next step, one confirmation on anything signed. Every validation ends in a verdict: accept, reject, or escalate. Never a feeling. Students record which checks they reached for, so the profile can say which ones they never have.

The student writes the concept and owns four decisions: what is optimised, what is determined, what is required, what is assumed. Translation into symbols and code is the AI's job. Those four are not.

Assessments written for one student's work

Every student's project goes a different direction, so no two students can usefully be asked the same question about it. The in-class assessment is written for the project in front of it.

The project runs against a client: a simulated director of operations with no technical training, who releases data only when a student's request is clear enough to earn it. Each request is judged in a fresh agent, so a dozen judgments are independent rather than one drifting conversation, and the client answers on a timer twice a day. A student who asks vaguely gets a delay, not a refusal — which is the lesson.

Review performance, homework performance and project performance all feed the same profile, and students can add to it themselves.

What scales, and what does not

Named separately, because they are different questions.

ScalesMeasured
Assessing a full round of reviewsunder three minutes for 50 students, as currently built
Scanning in-class papers and extracting the text to update profilesunder a minute at scale
Distributing student materials through Gitunchanged as class size grows

What does not scale is a human reading every review every day. At 50 undergraduates the review runs on completion credit rather than assessment — but completion credit does not mean no feedback. A student still learns they got 3 of 10, and it still updates their profile.

Running it in your course

In the dependency graph of this course, only two nodes are course content. The topology, the gates and the skills are the same whatever the subject.

Transfers unchangedYours alone
The router and the naming ruleThe lecture content
The four pipelines and their fixed orderThe problems, and what a correct answer in your field has to satisfy
The publishing gate that refuses unvalidated materialThe client dossier, if you run a consulting project
The ownership boundary, as a checkable file listYour students
One private repository per student

The skills are structurally generic. The worked examples inside them are logistics engineering, and a machine cannot carry a worked example across into molecular genetics and produce anything but confident nonsense. Those places are flagged rather than translated, and rewriting them is the actual work of adopting this.

Three of the disciplines are not about logistics at all. One asks whether a thing is tested. One asks what it costs per student at scale. One asks whether feedback is short enough for a student to act on.

Who owns what

Two sentences settle nearly every day-to-day question: the instructor owns what the course teaches; the mechanism owns how it runs. The asymmetry is deliberate, and it was the instructor's own request — not to become the bottleneck. The boundary is a file list rather than a permission, so it can be checked instead of remembered: a script exits 0 for ours, 1 for the instructor's, 3 for ask them.

Where it stands, honestly

Six students is a pilot, not a result. The student half runs every week and it works. Nobody outside NC State has run the staff half in another course. You would be the first, and we would rather say that now than have you find it out in week three.

The repositoryNot public yet. It is being repaired against its own validation pass before it opens. When it opens, it opens here.
JuliaA prerequisite today, not an option. The review skill opens by invoking it.
WindowsAddressed, not tested. Every read, write and subprocess decode passes an explicit encoding, so the failure that stops a pipeline on the first em dash is gone. Nobody has run the course on Windows.
SupportThere is none. Two people teach one course. The honest answer to a wall is to email us, and that is a conversation we want.

We run a validation pass on the toolkit itself, pointed at the artifact the way someone receiving it would read it rather than at the code that produced it. The most recent pass returned a list of things to fix before the repository opens, and that list is being worked through.

Talk to us

The reason this page exists is to find a course in another subject willing to try this. If you teach something where students could meet the concepts on their own and bring the difficulty with them, we would like to hear from you — including if your reaction is that it would not work in your field, and why.

We cannot promise a response time. We can promise a real answer.

Related