A high-throughput and memory-efficient inference and serving engine for LLMs.
See also[edit]
- AI, Inference, Agents, vLLM, Google TPU Ironwood (2025)
- LLM, MLLM, LoRA, LLaMA, LLaMA3, QLoRA, Falcon, PaLM 2, Gemini, Mixtral 8x7B, BitNet, Measuring Massive Multitask Language Understanding (MMLU), NVLM, LangChain, LlamaIndex, MCP, vLLM, DeepEval, NotebookLM, Costory.io