12 modules
Models
One import per task. Each page has a live demo, the code to copy, measured timings, and the base model with its license.
Audio models
01Image models
05- 01
Scan
OCRRead text from any photo. Pick the size.Download~31 MBFastest~0.6 s - 02
Remover
Background removalDrop a photo, get a transparent PNG.Download~219 MBFastest~15 s - 04
Alt
BetaImage captioningOne sentence for any image, in nine languages.Download~316 MBFastest0.7–1.2 s - 05
Depth
BetaDepth estimationA depth map from a single photo.Download~50 MBFastest0.8 s · 640×480 - 06
Spot
BetaObject detectionBoxes and labels for 80 everyday objects.Download~29 MBFastest~1 s
Text models
06- 07
Vector
BetaText embeddingsSemantic search in 23 MB.Download~45 MBFastest14 ms / sentence - 08
Lingo
BetaTranslationTranslate text without a server.Download~22 MBFastest~0.1 s / sentence - 09
Polish
BetaTranscript cleanupTurn raw dictation into written text.Download~339 MBFastest~1–2 s / sentence - 10
Emojify
BetaText-to-EmojiTurn a sentence into emojis. 4 MB model.Download~4 MBFastest10 ms / sentence - 11
Voice
BetaText-to-SpeechNatural speech, generated locally.Download~326 MBFastest0.7 s / sentence - 12
Imagine
BetaImage generationA picture from a sentence. On your GPU.Download~3.4 GBFastest~10–40 s · 512²