Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Beyond One-Size-Fits-All: Building Intelligent LLM Selection Systems
Learn how to deploy an OpenAI‑compatible LLM router that classifies prompts, selects the appropriate model, and optimizes cost, speed, and accuracy.
The presentation will demonstrate NVIDIA’s LLM Router Blueprint - a production-ready solution that intelligently routes user prompts to the most appropriate Large Language Model based on task classification and complexity analysis.
We’ll explore how this OpenAI-compatible routing system solves the common dilemma of choosing between accuracy, speed, and cost in modern AI applications. The session will cover:
Live Demo: Watch the router automatically classify prompts (code generation → DeepSeek, general Q&A → Llama 70B, simple rewrites → Llama 8B) and route them accordingly
Architecture Deep Dive: Understanding the three-component system (Router Controller, Router Server with Triton, Downstream LLMs)
Real-World Impact: Cost optimization strategies and performance benchmarks
Hands-On Implementation: Step-by-step deployment using Docker Compose and Kubernetes
Customization Possibilities: Creating custom routing policies for domain-specific use cases
Attendees will leave with practical knowledge of implementing intelligent LLM routing in their own applications, complete with code examples and deployment configurations.
NVIDIA's LLM router intelligently directs requests using Rust and Triton.
Compose Email
Loading recent emails...