Engineering breakdowns, technical guides, speed benchmarks, and announcements straight from our core infrastructure team.
Announcing the addition of FLUX.1 schnell and FLUX.1 dev models. Create photorealistic images with perfect text rendering in seconds.
Discover the architectural changes, custom CUDA kernels, and tensor parallel setups we implemented to make Llama 3 70B run blazingly fast.