Bonsai Image
State-of-the-art image generation, in your browser. Bonsai Image 4B is a compressed text-to-image model from PrismML, built for local generation on iPhone, Mac, and GPUs.
Try This Model NowLocateAnything
Detect and label objects in images and videos. LocateAnything is an NVIDIA vision-language model that finds objects, text, GUI elements, and points in images with natural language prompts.
Try This Model NowWhisper AI
Whisper AI is OpenAI’s speech recognition model for transcribing, translating, and understanding spoken audio.
Try This Model NowDeepSeek OCR 2
DeepSeek OCR 2 is an open-source OCR and document understanding model built for complex layouts, Markdown output, and human-like reading order.
Try This Model NowDeepSeek OCR
DeepSeek OCR is an open-source vision-language OCR model that converts document images into structured text and Markdown with efficient visual token compression.
Try This Model NowSCI Bot: what Sci-Hub's AI research assistant does and how to use it
Sci Bot is an AI-powered research assistant connected with Sci-Hub. It lets users ask scientific questions in natural language, then tries to answer with information drawn from research papers.
LTX-2
LTX-2 is an open-source AI video model that generates synchronized video and audio for creative, research, and production workflows.
Try This Model NowVoxCPM
VoxCPM is an open-source TTS model family for multilingual speech generation, voice design, and realistic voice cloning.
Try This Model NowIndexTTS 2
Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Try This Model Now