Open sourceActive

QuiLLMan

Voice chat example app built on Modal using Kyutai Lab's Moshi speech-to-speech model.

Open source page

Product features

Product features

Addenda

Configuring vLLM for maximum throughput

Deploying vLLM on Modal

FastAPI Server

Serves both frontend and backend, exposed on Modal via @app.asgi_app().

Loading filings from the SEC EDGAR Feed

Moshi Websocket Server

Loads a Moshi model instance and maintains a bidirectional websocket connection with the client.

Opus Audio Compression

Compresses audio across the network using the Opus codec for efficient transmission.

Organizing a batch job on Modal

React Frontend

Static React app using Web Audio API to record microphone audio and play back model responses.

Run LLM inference at maximum throughput

Serving tokens at maximum throughput

Transforming SEC filings for batch processing

Utilities for loading filings from the SEC EDGAR Feed

Utilities for transforming SEC Filings