Chalk is an AI lecture-capture platform, live at chalkrecap.com. Record a class, upload a file, or paste a video link, and it hands back a clickable table of contents, study notes, per-chapter quizzes, spaced-repetition flashcards, and a tutor you can ask questions. It launches from Canvas and Moodle, imports straight out of Zoom and Webex, and ships an Android app alongside the web. Under the hood it is a serverless media pipeline that survives being killed mid-job, which is a fun property to need.
Chalk is a monorepo with a Next.js 16 web app that carries the entire backend (82 API routes) and an Expo app that is a pure client of it. Storage is Neon serverless Postgres plus Vercel Blob for media, auth is a custom JWT system, billing is Stripe, and transcription runs on Groq's Whisper with OpenAI as the fallback.
The core promise: a two-hour lecture goes in, and a navigable library item comes out, with chapters that jump the video to the right second. Around that core sit the things a real course needs: binders that pull several documents into one study set, LMS launch from Canvas and Moodle, meeting imports from Zoom and Webex, and engagement analytics that show an instructor where the class actually stumbled.
Lecture recordings are where studying goes to die. A two-hour video with no structure means scrubbing around hoping to spot the whiteboard changing. Notes apps do not know what was said, and transcription tools give you a wall of text with no map.
Students need the lecture broken into topics, summarized, and quizzable, without doing any of that work themselves. That is a pipeline problem, not a note-taking problem.
Every lecture takes one of three doors in. Uploads go straight from the browser to Vercel Blob using short-lived tokens, because serverless request bodies cap out long before video sizes do. Live recordings stream in as segments that get stitched into one file. YouTube links never download video at all: Chalk pulls the video's own caption track and skips transcription entirely, then plays the video through YouTube's embed.
Then the orchestrator takes over: extract mono 16 kHz audio with ffmpeg, transcribe with Whisper, and hand the transcript to gpt-4o to segment into chapters. Processing runs inside Next.js after(), which keeps working after the response is sent, and an atomic database claim stops two instances from processing the same lecture twice.
Neon's serverless Postgres driver speaks HTTP, which fits functions that appear and vanish constantly. The schema lives in code as a versioned migration array with idempotent DDL, because two racing instances both have to be safe to run the bootstrap. Vercel Blob stores raw media, stitched recordings, exported clips, and thumbnails.
ffmpeg ships as a static binary and gets spawned directly as a CLI, no wrapper library, which means fewer moving parts and no binary-path weirdness on Windows. Auth is deliberately homegrown but small: bcrypt password hashing, HS256 JWTs signed with jose, delivered as an HttpOnly cookie for web and a bearer header for mobile, with a token version claim so plan changes and forced sign-outs propagate within about two minutes.
Chalk started on two pinned OpenAI models and got expensive, so the model layer became a routing decision instead of a constant. Transcription is the single biggest line in the bill, because it runs on every minute of every upload and cannot be skipped. Groq serves the same Whisper weights behind an OpenAI-compatible API at roughly a tenth of the price and, critically, still returns per-segment timestamps, which the entire pipeline is built on. OpenAI's own cheaper transcription model was disqualified for exactly that reason: it emits no segment timestamps at all. Without a Groq key everything falls back to whisper-1, so the cheap path was safe to deploy before the key existed.
Language work splits the same way, by whether the task is a judgment call or mechanical. Chaptering, transcript refinement, and translation are structural: find the topic shifts, copy a verbatim quote, emit the same number of lines back. Those run on a cheap tier at six to eight times less. The calls a user actually judges the product on, ask-the-video, the tutor, study packs, and deep dives, stay on the strong model. Every model call is metered at one chokepoint in the OpenAI client wrapper, so the usage ledger cannot drift as features get added, and a global spend circuit breaker can stop the expensive paths cold.
Two whole subsystems exist because language models are unreliable narrators. First, timestamps: gpt-4o drifts up to 100 seconds when asked for chapter start times, so Chalk instead asks it to return a verbatim quote from where the chapter begins, then finds that quote in the real transcript to resolve the true time. Second, coverage: the model nondeterministically stops chaptering partway through, so a refill loop re-asks for the uncovered tail up to four times.
Quizzes get a two-pass treatment to kill the AI-quiz smell: a first pass drafts questions in a professor voice with misconception-based wrong answers, then a second pass acts as an exam editor and rewrites anything lazy or guessable. There is also a per-chapter deep dive, a whole-lecture quiz, and an ask-the-video feature that answers free-form questions with clickable timestamp citations.
A table of contents is where studying starts, not where it ends, so the product grew a study layer. Study-pack flashcards run on an SM-2-lite spaced-repetition scheduler with four grades, and the scheduling math is pure and dependency-free precisely so the server and the client can share one source of truth: the server persists the next state, and the grade buttons show you the real next interval before you pick. Review state is per viewer, so a signed-in student and an anonymous browser each keep their own schedule.
Binders extend the same idea past a single video: pull several documents into one study set, with PDF and DOCX extraction, a cheap-tier cleanup and OCR pass for scanned pages, then per-document digests, a combined study guide, quizzes, and ask across the whole set. There is also a transcript translation mode for lectures delivered in a language the reader does not follow, built on the same structural contract as refinement, where line count and timestamps are untouchable because seeking, auto-follow, and search all hang off them.
For instructors there is engagement analytics: what students actually watched and which questions the class kept missing. The read model is owner-only and deliberately excludes the owner's own passes, so a teacher previewing their own lecture never skews the class picture. The waveform on the player used to be decorative, a hash of the lecture id shaped into sine curves, identical for an hour of speech or an hour of silence. It now measures real loudness second by second with ffmpeg, so you can see the quiet stretches before scrubbing into them.
Students do not start their day at chalkrecap.com, they start in an LMS, so Chalk implements LTI 1.3 properly: a JWKS endpoint, signed launch, and deep linking, so an instructor can drop a Chalk lecture into Canvas or Moodle as a normal course item. Lectures also arrive from Zoom and Webex through full OAuth connections with webhooks, so a recorded class can flow in without anyone re-uploading it.
Every one of those import paths is a retry risk, because webhooks fire more than once and OAuth callbacks get replayed. Imports are keyed idempotently against a unique constraint, so a duplicate delivery resolves to the lecture that already exists instead of paying to transcribe the same meeting twice.
The browser mints an upload token, pushes the video straight to Blob, creates the lecture row, and pokes the process route. The pipeline downloads to temp storage, checks the duration cap, meters the user's transcription minutes, extracts audio, and transcribes. Files over ten minutes are split into chunks that transcribe four at a time, and every finished chunk is checkpointed to the database.
That checkpointing matters because serverless functions get 300 seconds. If a run dies mid-transcription, the sweeper restarts it and it resumes from the saved chunks instead of paying Whisper twice. Lectures over 90 minutes get map-reduce chaptering in 30-minute windows. When segmentation lands, the lecture flips to ready, an overview generates from the chapter outline, and the viewer shows the clickable TOC.
Ownership checks live server-side in every route, including anonymous demo lectures that are scoped to a per-browser id. Admin routes sit behind a role gate. Rate limiting is a fixed-window limiter in Postgres that fails open on database trouble, backed by hard ceilings: 2 GB uploads, 4 hour media, 30 minute caps for anonymous users.
Whisper is the expensive step, so it is metered like a utility: free accounts get 300 transcription minutes a month, Pro gets 1500, and top-up packs exist for heavy semesters. Charging is idempotent per lecture, so retries never double-bill. Stripe handles the Pro subscription and one-time credit purchases through a signature-verified webhook that is the only code allowed to grant Pro.
The 300-second function ceiling shaped almost everything: the after() processing model, per-chunk checkpoints, claim locks with timeouts, and a daily cron sweeper that revives anything silent for six minutes and gives up after three attempts. There is no queue service because the checkpoints plus the sweeper turned out to be enough, and that is one less thing to run.
YouTube added its own drama by bot-checking datacenter IPs, so the importer requests captions the way a phone would, and the mobile app can even fetch captions client-side on its residential IP and hand them to the server. Chalk also has strong opinions about prose: every model response is banned from using em dashes, enforced in the prompt and then scrubbed with a regex, belt and suspenders. Raw transcripts are exempt because those are the speaker's words, not ours.
Chalk is live at chalkrecap.com, running on Vercel with the web app and an Expo client sharing one backend of 82 API routes. Android is in beta and iOS is next. Intake comes from recording, upload, video links, Zoom, and Webex; distribution goes back out through Canvas and Moodle over LTI 1.3. On top sit chapters, notes, quizzes, spaced-repetition study packs, binders, deep dives, clip export, translation, and ask-the-video.
The pipeline holds up on real inputs: four-hour lectures chunk, checkpoint, and chapter without babysitting, and the accounting stays correct even when the platform kills the process mid-job. The cost work matters as much as the reliability work, since a study tool that loses money on every upload is not a product. Routing transcription to Groq and the mechanical language work to a cheap tier cut the dominant per-hour costs by roughly an order of magnitude and six to eight times respectively, without touching the calls users judge the product on.