Fast serving framework for LLMs and vision models. Best for high-performance serving with structured generation.