ModelMesh and KServe bring eXtreme scale standardized model inferencing on Kubernetes
10/12/2021
This post announces the open-sourcing of ModelMesh, a model serving management layer for Watson products, and its integration with KServe. It details ModelMesh's architecture, core components (ModelMesh Serving, ModelMesh containers, runtime adapters), and supported model-serving runtimes (Triton Inference Server, Seldon MLServer). Key features highlighted include cache management and HA, intelligent placement and loading, resiliency, operational simplicity, and scalability demonstrated by packing 20K models into two serving runtime pods with single-digit millisecond latency. The post also announces the transition of KFServing to KServe and the collaborative development of a unified KServe API.