a bun cli · mit

Meeting recordings in,
minutes out.

Point momgen at an audio or video file and it writes the Minutes of Meeting. ElevenLabs Scribe does the transcript, a chat model of your choosing does the minutes, and the silence is cut out first so you never pay to transcribe it.

$ bunx momgen ./meeting.mp4
or install it — bun add -g momgen
  • ✓requires bun ≥ 1.2
  • ✓requires ffmpeg
  • ✓resumes after an interrupt
  • ✓caches every segment
one run, start to finish
recording
momgen ./standup-2026-08-04.mp4

And this is what lands in output/

Two files per run, in a fresh timestamped directory: transcript.md as Scribe heard it, and mom.md as your model summarised it. Plain markdown — paste it wherever the notes normally go.

mom.md1,284 tokens

Minutes of Meeting

2026-08-04 · 58m · standup-2026-08-04.mp4

Attendees

Sara (engineering), Amir (engineering), Nadia (design), Kian (product)

Decisions

  • —Ship the upload rewrite behind a flag, default off, until the load test clears.
  • —Design system migration moves to Q4; the shim stays in place until then.

Action Items

  • ☐Amir — load-test the new upload path(Aug 8)
  • ☐Sara — flag plumbing + rollback note(Aug 6)
  • ☐Kian — tell support about the Q4 slip(Aug 5)

Discussion Summary

Most of the hour went to the upload pipeline. Amir walked through the retry behaviour on partial uploads; Sara raised that the current path retries the whole file rather than the failed chunk, which is what the rewrite fixes. Nadia flagged that the design system migration would collide with the same release window…

how it works

Five stages, and the reason each one is fussier than it looks.

  1. 01

    Audio extraction

    ffmpeg strips the audio out of the video, keyed by the file's SHA-256 so a second run skips the decode entirely. The audio is re-timed on the way out, because call recordings tend to carry jittery timestamps that the mp3 muxer would otherwise drop speech over.

    If the extracted audio comes up materially shorter than the video claims — the signature of an interrupted download, whose header still promises the full length — the run warns instead of quietly transcribing a fraction of the meeting.

  2. 02

    Silence removal

    Transcription is billed per second of audio uploaded, and silence transcribes to nothing. Gaps longer than a second are cut before anything is sent, with a short pause left in place of each one so the words on either side don't run together and come back as a single mangled token.

    If the filter eats almost the whole recording — a wrong threshold for an unusually quiet source — the original audio is used instead. Paying to transcribe a mangled file is worse than paying for the silence.

  3. 03

    Cost estimate

    The speech-only audio is split into segments, the segments already in the cache are subtracted, and the run prints exactly what is left to upload with its price before asking whether to continue.

    Nothing reaches the provider until that prompt is answered. Decline it and no bytes leave the machine.

  4. 04

    Transcription

    Each uncached segment goes to ElevenLabs Scribe, and the full response is written to the cache before anything else runs — so an interrupted run resumes where it stopped and never pays for the same audio twice.

    Scribe is left to detect the language itself, which it does well on mixed Persian/English speech. Forcing a language code makes the loanwords worse.

  5. 05

    Minutes

    The full transcript is streamed to the chat model you configured, which writes the minutes: Attendees, Agenda, Decisions, Action Items, Discussion Summary.

    Streaming is not cosmetic. A reasoning model can think for minutes before its first token, and a non-streamed request sits idle long enough to be killed in transit. The run aborts only after five minutes with no delta at all — reasoning counts as progress.

bring your own model

Any endpoint that speaks /chat/completions.

The minutes model is reached through the AI SDK over a plain OpenAI-compatible endpoint, so the provider is two environment variables rather than a code change. Point LLM_BASE_URL at it, name the model in LLM_MODEL, done.

Transcription is ElevenLabs Scribe, and that part is not swappable.

.env
LLM_BASE_URL=https://opencode.ai/zen/go/v1
LLM_MODEL=deepseek-v4-flash
LLM_API_KEY=sk-…

The default. Nothing to configure.

configuration

Every knob, in one schema.

Nothing reads process.env outside a single module. A variable that is unset — or set to an empty string, which a half-filled .env line produces — falls back to its default. One set to nonsense fails at startup naming the variable, rather than surfacing as a $NaN cost estimate several minutes later.

Env varDefaultPurpose
LLM_API_KEY—required
ELEVENLABS_API_KEY—required
LLM_BASE_URLhttps://opencode.ai/zen/go/v1OpenAI-compatible endpoint, without /chat/completions
LLM_MODELdeepseek-v4-flashminutes model
ASR_LANGUAGEunset (auto-detect)language hint for Scribe
ASR_PRICE_PER_HOUR0.22rate used for the estimate
ELEVENLABS_STT_MODELscribe_v1Scribe model
SILENCE_THRESHOLD-35dBbelow this counts as silence
SILENCE_MIN_SECONDS1shorter gaps are left alone
SEGMENT_SECONDS900audio segment length
MOM_LANGUAGEEnglishlanguage of the minutes

Caches live in $TMPDIR/momgen-cache — extracted audio, speech-only audio, segments, and Scribe responses, keyed by model and language hint.

Stop writing up the meeting you just sat through.

$ bunx momgen ./meeting.mp4
or install it — bun add -g momgen