AGENT OBSERVATORY
NLโSQL ROUTER โข LIVE PROD โข REAL-TIME RAG TELEMETRY
yourdomain.com
โ
AGENT OBSERVATORY
NLโSQL ROUTER โข LIVE PROD
LIVE updated 14:43:00
TOTAL REQUESTS
1,284
in 7d range
ACTIVE SESSIONS
3
in-flight now
CONCURRENCY (PEAK)
12
max in-flight, range
AVG RESPONSE TIME
0.42s
seconds
P95 RESPONSE TIME
1.85s
seconds
ERROR RATE
0.0%
non-2xx / total
GPU UTILIZATION
78%
4x H100 SXM5
BACKEND SPLIT
85% / 15%
vLLM / Qdrant
Response Time โ p50 / p95 / p99
seconds, per interval
p50
p95
p99
Requests / sec by backend
vLLM vs Qdrant
vLLM Engine
Qdrant DB
Active Sessions / Concurrency
in-flight requests over time
in-flight (total)
System & H100 VRAM Allocation
host + GPU utilization
GPU 0-3 VRAM
Host CPU
โ๏ธ Host Setup & System Hardware
Validated host CPU architecture, server type, and GPU slot enumeration.
NVIDIA_GENERIC
Server Type
PASS
CPU Validation
4 Slots
GPU Slots
320 GB
Total VRAM
๐ Local Auth Mode
Automatic Admin Bootstrap
When local authentication is selected, an administrator account is automatically created on deployment:
Username / Email:
admin@yourdomain.comInitial Password:
Change_me123Role & Privileges: SUPERADMIN / OWNER
๐งฉ GPU-Aware Model Auto-Fitting Engine
Evaluates precision quantization or lower weight variants to fit available GPU hardware.
| Status | Model Identifier | Quant / Precision | Req VRAM | TP Size | Recommendation / Reason |
|---|
๐ฅ Ingest Enterprise Document
Chunk text and index high-dimensional embeddings into Qdrant vector storage.
๐ Indexed Qdrant Documents
Active vector collections in `excel_rag` Postgres & Qdrant database.
โก vLLM Multi-GPU Tensor Parallelism Status
Qwen2.5-72B
Active Model
4 GPUs
Tensor Parallel Size
26.4 ms
Avg TTFT
118.5 / s
Tokens / Sec / Stream
๐งช Launch Multi-User Load Test
Evaluate concurrent stream processing speeds across all 4x H100 GPUs.
๐ Benchmark Results
Click "Execute Async Benchmark" to measure throughput, TTFT percentiles, and TPS.
0Requests / Sec
0Cluster Tokens / Sec
0 msTTFT (P50)
0 sLatency (P90)