ClaudeD is a web app for musicians: upload a recording or paste a YouTube link, and it separates the mix into stems (vocals, drums, bass, guitar, piano), transcribes them toward sheet music, and opens an in-browser notation editor with a per-job AI chat assistant alongside it. It’s live at claudedtemp.com, self-hosted and deployed via a Cloudflare tunnel.
Pipeline
Separation runs on Demucs plus RoFormer/MDX-Net; transcription runs Basic Pitch through music21 into notation rendered with OpenSheetMusicDisplay. A newer “Photo to Score” tab does optical music recognition — point a camera at printed or handwritten sheet music and get back editable notation — via a three-engine chain (Audiveris → homr → Claude vision) that tries the highest-ranked engine first and falls back automatically. Scores can also carry an AI-generated decorative watermark, composited under the notation and served from PDF downloads.
The stack is FastAPI, Celery, and Redis on the backend, PostgreSQL via
SQLAlchemy/Alembic for storage, and a React/TypeScript/Vite frontend styled
with Tailwind — all wired together with Docker Compose and GitHub Actions
CI, with the Claude API doing the chat and vision-OMR work. A Capacitor iOS
shell wraps the same web app in a native WKWebView, device-verified with
camera capture, push notifications, haptics, and deep links.
Practice Studio
Past the pipeline is a practice layer: tap-along and play-along modes that follow a moving cursor through the score, a subsection-practice mode for looping a bar range, and a flashcard mode for memorization drilling.
An engineering-decision story
For a while the notation renderer was a three-way split — OpenSheetMusicDisplay in most places, Verovio on a couple of newer surfaces, and a paid Flat.io embed for the in-browser score editor. Each had different cursor behavior and a different WASM/JS payload. Rather than keep patching around the seams, I ran OSMD, Verovio, and Flat.io side by side against the same real scores, picked OSMD as the single engraver everywhere, and tore the other two out. The production bundle went from 11.66 MB to 4.07 MB — a 65% cut — with the practice surfaces staying behaviorally identical throughout the swap.
A real ML research track
Running alongside the product is a from-scratch flute audio-to-notation model: a pretrained audio encoder feeding a small note-event head, a beat-aware quantizer, and a seq2seq notation-cleanup pass, evaluated against a frozen real-audio holdout that’s never trained on. The research track keeps session-by-session experiments alongside internal benchmarking that ranks the OMR engines against each other on real scanned/photographed scores.
Where it stands
Automated backend and frontend suites gate changes. Feature plans, including the OSMD consolidation above, are planned and executed through the same AI-assisted planning system described in the AI-native workspace write-up.