TLC — Teacher's Lesson Creator

A teacher says what the lesson is about, the grade level, how long the period runs, and whether they want homework and a demonstration. Two specialised model adapters take it from there, one owning structure and one owning engagement, and their two drafts are merged by fixed rules into a complete lesson: objectives, a sequence, an assessment, an answer key and the materials list.

Glover Works' own project, built for the Gemma 4 Good Hackathon. No client involved.

Planning a lesson well takes longer than teaching it. TLC takes the decisions a teacher has already made and turns them into a lesson they can run.

It was built for the Gemma 4 Good Hackathon, run by Kaggle and Google DeepMind, which required a working demo, a public repository and a technical write-up. All three are linked at the end of this page.

Challenge

A general-purpose model will write something that looks like a lesson plan. Getting one that a teacher can actually run is a different problem, and it has three parts.

  • The output has to be the same shape every time, because the application has to render it and a teacher has to scan it.
  • It has to be right. A lesson is not improved by a confidently wrong definition.
  • The two things a lesson needs — rigor and engagement — pull against each other, and asking one model for both tends to produce a compromise that is neither.

Approach

Two adapters rather than one model and a longer prompt, and both of them work on every phase rather than one drafting and the other editing. Hunter owns structure and rigor: the learning objective, the lesson sequence, the assessment, the answer key, the time arithmetic and standards alignment. Christine owns depth and engagement.

They build in parallel, so neither is anchored to the other's first idea. A single review pass then looks at both scaffolds together and flags where they disagree, and both write again with that review in front of them.

Only then are the two combined, by deterministic field-ownership rules in ordinary code rather than by asking a model to merge them. The same inputs produce the same lesson, and a missing required field throws rather than being quietly papered over.

Both were trained with QLoRA on google/gemma-4-e4b-it — supervised fine-tuning over an NF4 base with bf16 compute, LoRA rank 16, alpha 32, dropout 0.05. The authors describe the training data as roughly 250 schema-validated outputs generated by Gemma 4 31B acting as a teacher model across a K-12 topic and grade matrix.

Solution

What the system does, in order:

  • A teacher gives the topic, the grade, the length of the period and whether they want homework and a demonstration, and can upload their own source material.
  • Both adapters write a scaffold, in parallel, as strict JSON through the model's native function-calling.
  • A review pass reads both scaffolds together and reports where they disagree.
  • Both write a final package, in parallel, with that review in front of them.
  • The two packages are merged by field ownership, with materials and homework unioned and de-duplicated by name.

Around all of that: every generation is validated against a schema, and a failure is retried once with the error appended rather than accepted and rendered. Vocabulary and misconception claims are checked against the Wikipedia and Wikidata APIs, and a contradiction sends the lesson back. Every piece of content carries a label saying whether it traces to the teacher's own source, was shaped by it, or was generated outright.

On where it runs, precisely: the adapters are the published local-inference path, converted to GGUF and served through llama.cpp with both loaded so the worker swaps between them per request. The hosted demo runs against cloud Gemma 4 31B.

Outcome

Verified outcomes currently available:

  • A working demo, a public MIT-licensed repository and a technical write-up, which is what the competition required.
  • Two MIT-licensed adapters published with their training method, data description and hyperparameters.
  • Schema-validated output with a documented retry, and a merge step whose failure mode is an exception rather than an incomplete lesson.

What this demonstrates

Model adaptation: QLoRA training against a published base model, with the method and hyperparameters public.

Specialised model roles: two adapters with different jobs, rather than one model asked to do everything at once.

Structured generation: schema-enforced output with a retry, because an application cannot render prose that was supposed to be JSON.

Deterministic assembly: the part that combines two model outputs is ordinary code with fixed rules, which is what makes the result reproducible.

Local inference: a published path that runs on hardware an organization controls, for material that cannot leave the building.

A similar problem?

Most of our work starts with someone describing a mess. Describe yours.

Start a conversation