BentoML
Framework for building production-ready AI application services.
About
BentoML turns model inference code into production services. Developers define a service in Python with type-hinted APIs, run it locally, then package the model, code, and dependencies into a standardized artifact called a Bento that builds into a Docker image for deployment anywhere. The framework is agnostic to ML framework and modality, serving LLMs, embeddings, image generation, audio, and computer vision workloads. Performance features include adaptive request batching, worker parallelization, model composition, multi-stage inference pipelines, and GPU support with CUDA detection, while built-in observability covers monitoring out of the box. Beyond self-managed Docker or Kubernetes deployments, the company behind it offers BentoCloud as a managed platform with autoscaling and distributed serving. BentoML is open source under the Apache 2.0 license and is used by ML platform and backend teams that need reliable, cost-efficient inference APIs and multi-model serving systems.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- AI Deployment & MLOps
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- Apache-2.0
- Added
- Apr 3, 2026
Related Tools
Self-hosted Go gateway that routes LLM traffic across providers with failover, caching, and guardrails.
Open-source orchestrator for AI training and inference across clouds, Kubernetes, and bare metal.
Kubernetes-native workflow orchestration platform for machine learning and data pipelines.
Open-source AI gateway that routes requests to more than 1,600 LLMs through one API with guardrails and caching.
Container tool by Replicate for packaging ML models as standard Docker images.
Local AI API platform that runs LLMs on your hardware with OpenAI-compatible API.