✅ Customization Switch model ID routing at runtime with profiles Run concurrent models with a custom DSL swap matrix (#643)
Almost all configuration settings are optional and can be added one step at a time:
Advanced features matrix to run concurrent models with a custom swap logic DSL hooks to run things on startup macros reusable snippets
How does llama-swap work? When a request is made to an OpenAI compatible endpoint, llama-swap will extract the model value and load the appropriate server configuration to serve it. If the wrong upstream server is running, it will be replaced with the correct one. This is where the "swap" part comes in. The upstream server is automatically swapped to handle the request correctly.
In the most basic configuration llama-swap handles one model at a time. For more advanced use cases, using a matrix allows multiple models to be loaded at the same time. You have complete control over how your system resources are used.