The Proxy is the local scheduling and knowledge-injection component in CacheRoute. It receives requests forwarded by the Scheduler, selects a local Instance, prepares text or KVCache injection, and forwards the request to the selected Instance.
In the two-level CacheRoute scheduling pipeline, the Scheduler decides which LLM system / Proxy should receive a request, while the Proxy decides which local Instance should execute it and how knowledge injection should be performed.
Client
└──> Scheduler
└──> Proxy
├── maintains a local Instance pool
├── receives Instance resource snapshots
├── selects a local Instance
├── prepares text or KVCache injection
└── forwards the request to Instance / vLLM + LMCache
| Plane | Default port | Description |
|---|---|---|
| Service plane | 8001 |
Receives Scheduler-forwarded OpenAI-compatible requests. |
| Control plane | 8002 |
Receives Instance registration, heartbeat, topology reports, resource snapshots, and debug queries. |
| Browser UI | 8202 |
Displays Proxy health, Instance liveness, resources, topology, Scheduler registration, charts, filters, and raw diagnostics. |
The Proxy can run without a Scheduler during local demos. In that case, Scheduler registration may fail non-fatally, but Instance registration and resource reporting can still be validated through the Proxy control plane and Proxy UI.
proxy/
├── proxy.py # Proxy service plane and startup lifecycle
├── proxy_cli.py # Proxy CLI for status inspection
├── sclient/
│ └── scheduler_client.py # Proxy -> Scheduler control-plane client
├── resource/
│ ├── instance_pool.py # Instance pool and normalized resource state
│ ├── p_control_plane.py # Proxy control plane
│ └── hb_log.py # Heartbeat reporting
├── strategy/
│ ├── base.py # Base Instance selection strategy
│ ├── round_robin.py # Round-robin Instance selection
│ ├── least_load.py # Experimental least-load Instance selection
│ ├── least_inflight.py # Reserved for future strategy extension
│ └── factory.py # Strategy builder
├── queue/
│ ├── manager.py # Prepare/ready queue manager
│ ├── task.py # ProxyTask state
│ ├── instance_queues.py # Per-Instance queues
│ └── knowledge.py # Knowledge retrieval and injection helpers
└── README.md
The browser UI implementation lives outside this directory under UI/proxy_ui/.
Start Proxy from the test directory:
cd test
python3 demo_proxy.py \
--host 127.0.0.1 \
--port 8001 \
--strategy round_robin \
--injection-strategy iwsdemo_proxy.py starts the Proxy browser UI by default and prints:
[demo_proxy] Proxy UI available at: http://127.0.0.1:8202
Open:
http://127.0.0.1:8202
Common options:
| Option | Description |
|---|---|
--host |
Proxy service-plane bind host. The demo also uses it as the advertised host. |
--port |
Proxy service-plane bind port. The demo also uses it as the advertised port. |
--strategy |
Local Instance selection strategy. Defaults to round_robin; use least_load for the experimental load-aware selector. |
--injection-strategy |
Knowledge injection strategy. Use default or iws. |
--ready-release-policy |
Ready queue release policy: ordered or text_bypass. |
--kdn-links-json |
Optional static KDN topology metadata. |
--proxy-ui |
Explicitly enable the browser Proxy UI. This is already the default. |
--no-proxy-ui |
Disable the browser Proxy UI subprocess. |
--proxy-ui-listen HOST:PORT |
UI server listen address, default 127.0.0.1:8202. |
--proxy-ui-url URL |
Browser-facing URL printed in logs, useful for tunnels / forwarded ports. |
The Proxy UI is the main browser observability dashboard for the Proxy runtime. It is frontend-only and does not mutate Scheduler routing, Proxy Instance selection, injection strategy, KDN behavior, KVCache behavior, or Instance forwarding.
Default URL:
http://127.0.0.1:8202
It shows:
- Proxy health and
/debug/status. - Scheduler registration state, best effort.
- Instance pool and TTL-derived alive / stale state.
- Instance resource snapshots reported by demo Instances.
- Per-Instance resource cards and sortable tables.
- CPU, memory, GPU, network, and alive/stale trend charts.
- KDN topology links.
- Raw diagnostic JSON with copy actions.
- Local controls for refresh, pause/resume, polling interval, filters, search, sort, chart history, and theme.
The UI server proxies requests through UI/proxy_ui/proxy_ui_server.py so the browser does not need direct CORS access to Proxy or Scheduler APIs.
See UI/proxy_ui/README.md for details.
Proxy startup
├── initialize InstancePool
├── start Proxy control plane on :8002
├── start Proxy browser UI on :8202 in demo mode, unless disabled
├── load Instance selection strategy
├── register to Scheduler control plane, non-fatal for local demos
├── report topology metadata to Scheduler, if configured
└── start heartbeat loop
During shutdown, the Proxy tries to unregister from the Scheduler. If the process is killed directly, Scheduler removes it after heartbeat expiry. In demo mode, the UI subprocess started by demo_proxy.py is cleaned up on exit.
GET /healthz
GET /debug/status
POST /v1/instance/register
POST /v1/instance/heartbeat
POST /v1/instance/unregister
GET /v1/instance/list?include_dead=true
Instances register static information such as instance_id, host, port, endpoints, tags, weight, and meta. Heartbeats refresh last_seen_at and can optionally report lightweight non-lifecycle load fields such as qps_1m and gpu_util; Proxy request lifecycle hooks maintain inflight.
After PR #87, demo Instances can report host resource snapshots to the Proxy control plane:
POST /v1/instance/resource_snapshot
GET /debug/instance_resources
GET /debug/instance_loads
Use /debug/instance_loads to inspect Proxy-maintained per-Instance inflight counters, qps_1m when present, queue-pressure hints when the queue snapshot provider is available, and the queue-aware least_load score used for Instance selection.
curl -sS http://127.0.0.1:8002/debug/instance_loads | python3 -m json.toolleast_load uses a conservative deterministic score with three load-signal layers:
- Proxy lifecycle pressure:
inflight. - Instantaneous queue pressure:
active_prepare,active_ready,prepare_queue_depth, andready_queue_depth. - Predicted reservation pressure from QueueManager timelines:
pending_prefill_count,active_decode_count, andpredicted_total_backlog_ms.
load.inflight remains the primary signal with implicit weight 1.0. Queue hints add per-Instance active prepare work, active ready work, prepare queue depth, ready queue depth, pending prefill reservations, active decode reservations, and predicted backlog time. qps_1m from Instance heartbeats is supported but disabled by default with weight 0.0. Missing predicted fields remain unavailable and fall back to the instantaneous queue-aware behavior; missing queue hints fall back to inflight-only scoring. If no load signal is known for any candidate, selection falls back to round-robin. Ties at the minimum score are also broken by round-robin.
Default least_load weights are:
| Component | Default weight | Source |
|---|---|---|
inflight |
1.0 |
Proxy lifecycle counter |
active_prepare |
1.0 |
Proxy queue manager per-Instance snapshot |
active_ready |
1.0 |
Proxy queue manager per-Instance snapshot |
prepare_queue_depth |
0.25 |
Proxy queue manager per-Instance snapshot |
ready_queue_depth |
0.25 |
Proxy queue manager per-Instance snapshot |
pending_prefill_count |
1.0 |
Proxy QueueManager reservation snapshot |
active_decode_count |
0.5 |
Proxy QueueManager reservation snapshot |
predicted_total_backlog_ms |
0.001 |
Proxy QueueManager reservation snapshot |
qps_1m |
0.0 |
Instance heartbeat |
The debug response includes the raw pressure fields (pending_prefill_count, active_decode_count, next_slot_ready_in_ms, prefill_free_in_ms, and predicted_total_backlog_ms) plus least_load_score.total, raw least_load_score.components, weighted components, metric sources, queue source availability, and the weights used to compute the score.
The reporting path is:
test/demo_instance.py
├── starts or reuses Rust Resource Agent
├── waits for Resource Agent /healthz
├── starts reporting only after Instance registration succeeds
└── POSTs snapshots to Proxy /v1/instance/resource_snapshot
Successful snapshot updates are no longer logged on every report at INFO. The Proxy logs the first successful snapshot per Instance and uses debug-level logging for repeated successful updates.
The resource state appears in both APIs:
curl -sS "http://127.0.0.1:8002/debug/instance_resources" | python3 -m json.tool
curl -sS "http://127.0.0.1:8002/v1/instance/list?include_dead=true" | python3 -m json.toolNormalized resource fields include:
| Field | Meaning |
|---|---|
cpu_util |
CPU utilization percentage from the agent snapshot. |
memory_used_mb / memory_total_mb / memory_free_mb |
Host memory snapshot. |
memory_free_ratio |
Admission-oriented memory free ratio. |
gpu_util_avg |
Average GPU utilization if GPUs are visible. |
gpu_mem_used_mb / gpu_mem_total_mb |
Aggregated GPU memory. |
network_rx_mbps / network_tx_mbps |
First observed network interface throughput. |
admission_state |
Agent capacity hint, such as accepting, degraded, or rejecting. |
resource_ts_ms |
Agent collection timestamp. |
resource_reported_at |
Proxy receive time in seconds. |
resource_report_monotonic_ms |
Reporter monotonic timestamp. |
resource_report_wall_time_ms |
Reporter wall-clock timestamp. |
reported_instance_id |
Instance ID carried by the report metadata. |
raw_resource |
Raw agent snapshot retained for debugging. |
Resource snapshots are currently observability data. They are not yet used by the active Instance selection strategy.
POST /v1/topology/report
GET /v1/topology/kdn_links
Instances can report measured KDN link metrics. When multiple Instances report the same KDN, the Proxy keeps the best link using higher bandwidth and lower latency as the preference rule. The merged topology can later be reported to the Scheduler.
A Scheduler-forwarded request is processed as follows:
Scheduler
└──> Proxy service plane :8001
├── recover CacheRoute Request
├── build OpenAI-compatible Instance request body
├── select a local Instance
├── optionally run IWS injection decision
├── create ProxyTask
├── enqueue task into prepare queue
├── prepare text or KVCache injection
├── move task into ready queue
├── forward request to Instance
└── return Instance response to Scheduler
Supported service-plane endpoints:
POST /v1/chat/completions
POST /v1/completions
For streaming chat completion, the Proxy forwards the downstream SSE stream and appends a cacheroute_meta SSE event before [DONE]. This metadata contains CacheRoute timing and injection traces.
| Strategy | Aliases | Status | Description |
|---|---|---|---|
round_robin |
round_robin, round-robin, rr |
Default | Selects alive Instances in round-robin order. |
least_load |
least_load, least-load, ll |
Experimental | Selects the lowest known Instance load by load.inflight, then uses qps_1m as a secondary signal. Missing metrics remain unknown rather than zero; if all candidates lack usable load metrics, selection falls back to round-robin. |
kv_aware |
Planned | Planned | Future strategy intended to consider KVCache locality/inventory together with runtime load. |
Select least_load from the demo CLI:
python3 test/demo_proxy.py --strategy least_loadOr select it through the environment:
PROXY_INSTANCE_STRATEGY=least_load python3 test/demo_proxy.pyIf no alive Instance is available, the Proxy returns:
503 no_instance
The strategy interface is extensible. Future policies can use queue state, resource snapshots, prefix locality, KVCache inventory, or KDN topology.
| Mode | Description |
|---|---|
default |
Uses the injection mode carried by the Scheduler request. |
iws |
Dynamically selects text injection or KVCache injection based on predicted cost. |
The IWS mode estimates text-injection and KVCache-injection costs, then chooses KVCache only when the expected benefit is large enough.
Simplified cost shape:
text_total
= max(text_prepare_wait, ready_wait)
+ text_prefill_service
kvcache_total
= max(kvcache_prepare, ready_wait)
+ redis_load
+ residual_prefill
choose KVCache if:
kvcache_total + kdn_queue_penalty + decision_margin < text_total
This lets the Proxy avoid KVCache when the KDN path is congested and use KVCache when transfer latency can be hidden or prefill savings are large.
The Proxy uses two stages for each Instance.
The prepare queue handles knowledge preparation:
- fetch knowledge from KDN;
- classify knowledge into
kv_ready,text_only, andmiss; - inject retrieved text into the prompt;
- trigger KVCache injection through Instance control plane;
- collect timing traces.
Concurrency is controlled by:
PREPARE_CONCURRENCY
The ready queue controls when prepared tasks are forwarded to the selected Instance. It maintains a predicted execution timeline with slot readiness, prefill start, first token, decode tail estimate, predicted queue wait, and predicted TTFT.
Concurrency is controlled by:
READY_CONCURRENCY
| Policy | Description |
|---|---|
ordered |
Release tasks in prepare sequence order. |
text_bypass |
Allow text tasks to bypass blocked KVCache tasks within a configured limit. |
Relevant variables:
PROXY_READY_RELEASE_POLICY
PROXY_TEXT_BYPASS_MAX_PER_FLUSH
Proxy
├── fetch knowledge metadata from KDN
├── identify kv-ready knowledge IDs
├── estimate KDN-to-Instance KV transfer time
├── reserve KDN KV link
├── call Instance control plane POST /v1/kv/inject_ready
├── wait for KV injection acknowledgement
└── forward request to Instance
If KVCache injection fails or no KV-ready knowledge exists, the Proxy falls back to text-only behavior and records the fallback path in the task trace.
python3 proxy/proxy_cli.pyUseful commands:
| Command | Description |
|---|---|
:status |
Show Proxy control-plane health and Instance counts. |
:instances [N] |
List alive Instances. |
:instances --all [N] |
List all Instances, including expired ones. |
:watch [--all] [--interval S] [--limit N] |
Continuously refresh Proxy status. |
:scheduler |
Query Scheduler control plane. |
:exit / :quit |
Exit. |
The browser Proxy UI covers the same observability surface and adds charts, cards, filters, and raw diagnostic copy actions.
cd test
python3 demo_proxy.py \
--host 127.0.0.1 \
--port 8001 \
--strategy round_robin \
--injection-strategy iwsOpen the Proxy UI printed by the demo:
http://127.0.0.1:8202
cd test
python3 demo_instance.py \
--host 127.0.0.1 \
--port 9001 \
--proxy-cp-url http://127.0.0.1:8002curl -sS "http://127.0.0.1:8002/debug/instance_resources" | python3 -m json.toolOr use the browser UI:
http://127.0.0.1:8202
python3 test/demo_resource_monitor_e2e.py \
--agent-listen 127.0.0.1:19201 \
--agent-url http://127.0.0.1:19201Use a non-default agent port if another Resource Agent is already running.
Common environment variables:
| Variable | Description |
|---|---|
PROXY_ID |
Proxy ID reported to Scheduler. |
PROXY_ADVERTISE_HOST |
Host reported to Scheduler. |
PROXY_ADVERTISE_PORT |
Service-plane port reported to Scheduler. |
PROXY_CP_HOST / PROXY_CP_PORT |
Proxy control-plane bind address. |
PROXY_INSTANCE_STRATEGY |
Local Instance selection strategy. |
PROXY_INJECTION_STRATEGY |
Injection strategy, default or iws. |
PROXY_READY_RELEASE_POLICY |
Ready release policy. |
PROXY_KDN_LINKS_JSON |
Static KDN topology metadata. |
PROXY_INSTANCE_TTL_S |
Instance alive TTL. |
PREPARE_CONCURRENCY |
Prepare queue concurrency. |
READY_CONCURRENCY |
Ready worker concurrency. |
PROXY_UI_LISTEN |
Optional default for the Proxy UI listen address. |
PROXY_UI_URL |
Optional browser-facing URL printed by demo_proxy.py. |
- Resource snapshots are visible in Proxy state but do not yet drive routing.
- Repeated successful resource reports are intentionally quiet to avoid log flooding.
unknown_instancewarnings usually mean a stale or external Instance process is still heartbeating to the Proxy control plane.- The Proxy UI is the preferred visual entry point for Proxy observability during demos and experiments.