Skip to main content

Command Palette

Search for a command to run...

Why Self-Host Agentic RAG

Published
•3 min read•View as Markdown
A

Anton R Gordon, widely known as Tony, is an accomplished AI Architect with a proven track record of designing and deploying cutting-edge AI solutions that drive transformative outcomes for enterprises. With a strong background in AI, data engineering, and cloud technologies, Anton has led numerous projects that have left a lasting impact on organizations seeking to harness the power of artificial intelligence.

Agent teams (planner, retriever, ranker, synthesizer) thrive when latency is predictable, privacy is strict, and costs are controlled. That’s why I deploy them on Amazon EKS with vLLM as the serving layer: you keep data in-VPC, pin workloads to GPUs, and scale horizontally under bursty traffic. vLLM’s design — continuous batching, optimized kernels, and a novel memory model — was built for high-throughput serving.

Core Concepts You Must Know

KV Cache (Keys/Values)

Transformers store “past” attention tensors as a KV cache so each new token can attend without recomputing history. The cache grows with sequence length and batch size; it’s the dominant memory pressure in generation. Practical knobs include limiting concurrent sequences, capping max context, and enabling quantized KV (e.g., FP8/INT4) to trade a little precision for longer contexts and higher throughput.

Paged Attention

Traditional serving wants one large, contiguous KV region per request — wasteful and fragmentation-prone. Paged Attention treats KV like virtual memory: split into fixed-size blocks (“pages”) and pack them non-contiguously. The result is higher concurrency, fewer OOM events, and better GPU utilization at long context lengths — exactly what multi-agent topologies need.

Think of it this way…

Paged attention: treat memory like pages in a book. The server stores token history in small blocks and mixes them to fill the GPU tightly. This cuts waste, reduces out of memory errors, and lets more requests run at once with longer context. Tune page size and watch memory headroom.

KV cache: think of it as the running notes of the model. It grows with every token and with every active request, so it is the main source of GPU memory use. Keep context lengths reasonable, cap concurrent sequences, try lower precision cache, and track evictions to avoid slowdowns.

A Practical Multi-Agentic RAG Topology

  1. Planner & Tool Router: chooses retrieval hops (hybrid vector + keyword), tools, and model selection.

  2. Retriever Layer: BM25 + embeddings, optional re-ranking, and domain-aware chunking.

  3. vLLM Pools: one Deployment per model/quant for blast-radius control; continuous batching turns spiky agent traffic into steady GPU work.

  4. Sidecar LLM (optional): a lightweight LiteLLM proxy collocated with agents that exposes an OpenAI-style endpoint, provides request policies and cost controls, and can route to Bedrock, local vLLM, or other backends while your primary model stays self-hosted.

EKS Scheduling: Taints, Tolerations, and Affinity

Create a dedicated GPU node group and taint it (for example, gpu=true: NoSchedule). Only Pods with matching tolerations (your vLLM Deployments) will land there, protecting GPUs from general workloads. Add node affinity or nodeSelector to target specific GPU SKUs and isolate traffic tiers.

Checklist

  1. Set Pod resources.requests/limits to match GPU memory; align container memory to expected KV cache growth.

  2. Add tolerations for the GPU taint; use node affinity for the GPU group.

  3. Expose health probes (e.g., /health) and auto scale on request rate or tokens/sec, not CPU.

  4. Split read-heavy embeddings from generation-heavy LLM pools to avoid contention.

Throughput Levers That Matter

  1. Tune max concurrent sequences, chunked prefill, and scheduler batch size to keep GPUs saturated without thrashing memory.

  2. Prefer models and quantization's that play well with paged KV; enable quantized KV when latency and context length trump tiny accuracy deltas.

  3. Watch cache residency: long contexts plus many agents' balloon KV; paged blocks curb fragmentation, but you still need sane upper bounds.

Operational Notes

Instrument tokens/sec, queue depth, P50/P95 latency, and GPU memory headroom. Capture prompt/response sizes per agent to spot hot paths. Keep a small “canary” pool on separate nodes for rolling updates and use per-agent budgets to prevent runaway tool loops.

Why This Belongs in Your Job Description

As a modern Agentic AI Architect or Chief AI Architect, it’s no longer just about understanding how agents operate anymore. You must be fluent in agentic AI on Kubernetes/EKS: GPU scheduling with taints/tolerations and affinity, observability for throughput and cost, and the memory model (KV cache plus Paged Attention). Combine these with a sidecar gateway like LiteLLM for policy and portability, and multi-agentic RAG becomes fast, private, and economical but on your terms.

More from this blog

Untitled Publication

39 posts