Ben Blum

תיאור מקרה

Voice notes

A web client and a FastAPI service. The design stores a recording and returns Markdown, with quotas around transcription.

כתובת האתר החי היא מציין מקום עד שהאפליקציה תעלה לאוויר. זה לא אתר פעיל.

טכנולוגיות

  • Vite
  • React
  • FastAPI
  • Python
  • Supabase Auth
  • Supabase Storage
  • Groq Whisper
  • Render

צילומי מסך

Placeholder screenshot of the voice notes recorder. Not a capture of the running app.
Placeholder. Not a capture of the running app. Intended view: recorder.
Placeholder screenshot of a voice notes transcript. Not a capture of the running app.
Placeholder. Not a capture of the running app. Intended view: Markdown transcript.

Problem

A recorded note should come back as Markdown. A demo that calls a transcription API needs quotas so cost and abuse stay bounded.

Architecture

The design splits the product into a web client and a Python API.

  • The client is a Vite React app. It records audio and uploads it to a private Supabase Storage bucket, voice-audio, at {uid}/{note_id}.webm.
  • The API is FastAPI on Render (Docker). POST /notes takes an audio_path after that upload and returns 202 with an id. GET /notes/{id} returns status and transcript_md. GET /notes/{id}/markdown returns text/markdown. GET /healthz is the health check.
  • Every route except /healthz requires a Supabase Auth JWT, checked on the server against JWKS, including aud and exp.
  • A background task downloads the file, sends it to Groq whisper-large-v3, writes Markdown, and updates status. The notes table defaults language to he.
  • Per-user daily usage is stored in voice.usage.

Decisions

  • The API stays in Python. A rewrite to a Deno Edge Function was rejected because the existing service is already Python.
  • A public copy is a new repository built from the current tree, with one commit after gitleaks. The private repository stays private. The steps are in docs/publish.md.
  • Render is specified as one Docker web service on plan 0.5c-512mb, with numInstances 1 and no scaling block. Whether to start on the free instance, which sleeps after 15 minutes without traffic, is still an open decision. This portfolio does not deploy the service.
  • Quotas in the design: 10 minutes of audio per user per day, 10 MB and 5 minutes per file, 120 minutes per day across users, and 1,500 Groq minutes per month. Past a cap, the API returns 429. TRANSCRIPTION_ENABLED=false makes POST /notes return 503 and skip Groq.
  • Rate limits in the design: 10 requests per minute per user and 30 per minute per IP.
  • Logs in the design omit audio, transcript, email, and JWT. The user id is hashed.
  • Samples stay synthetic or text-to-speech. The design excludes Ben's voice until that recording is approved.
  • Open or invite-only access for a public demo is an open question.

Tests

The design calls for pytest with Groq mocked. A run against real audio is manual, and only after sample recordings are provided. The benchmark table stays a placeholder until then. The high-level design also notes that this stack has not yet been checked on real audio.

Sources

The benforcapita/voice-notes README was not available while this page was written (the GitHub API responded not found). The sections above follow the portfolio low-level design.

כל הפרויקטים