Eight MI300X GPUs, Six Open Models, Five Routing Objectives

The first AMD Developer Cloud deployment guide showed how to put vLLM Semantic
Router in front of a balance-oriented ROCm backend. The maintained
multi-objective recipe takes the next step: clients choose the optimization
objective they want, while the router keeps each objective's signals,
projections, decisions, algorithms, and plugins isolated.
This guide deploys six physical open models across seven serving GPUs, reserves the eighth GPU for router classifiers, and presents five stable Mixture-of-Models entrypoints. Requests move between checkpoints with different architectures, latency, tool-use, and quality profiles instead of simulating those differences with aliases on one backend.


