Model Serving
7 items tagged with "model-serving"
Stacks6
MLOps Stack
MLflow, Kubeflow, TensorFlow, Feature Store - ML lifecycle
MLflow MLOps Stack
End-to-end MLOps pattern using MLflow for experiment tracking, model registry, packaging, and deployment, integrated with feature, data, and serving layers.
Kubeflow ML Platform
Kubernetes-native ML platform: Kubeflow Pipelines, training operators, KServe serving, and Katib tuning run the ML lifecycle on Kubernetes.
vLLM + Ray Serve
A high-throughput LLM serving stack combining vLLM's optimized inference engine with Ray Serve for scalable, multi-replica deployment.
Amazon SageMaker MLOps
A managed MLOps stack on AWS using Amazon SageMaker to build, train, deploy, and monitor machine learning models end to end.
Triton Inference Server
A high-performance model-serving stack using NVIDIA Triton to serve models from any framework with GPU optimization on Kubernetes.