AGENT OBSERVATORY

NLโž”SQL ROUTER โ€ข LIVE PROD โ€ข REAL-TIME RAG TELEMETRY

yourdomain.com
โ—†

AGENT OBSERVATORY

NLโž”SQL ROUTER โ€ข LIVE PROD
LIVE updated 14:43:00
TOTAL REQUESTS
0 5K 10K
1,284 in 7d range
ACTIVE SESSIONS 3 in-flight now
CONCURRENCY (PEAK) 12 max in-flight, range
AVG RESPONSE TIME 0.42s seconds
P95 RESPONSE TIME 1.85s seconds
ERROR RATE 0.0% non-2xx / total
GPU UTILIZATION 78% 4x H100 SXM5
BACKEND SPLIT 85% / 15% vLLM / Qdrant
Response Time โ€” p50 / p95 / p99 seconds, per interval
p50 p95 p99
5.0 3.5 2.0 0.5 0 00:56 01:13 01:30 01:46 02:03 02:20 02:36
Requests / sec by backend vLLM vs Qdrant
vLLM Engine Qdrant DB
0.07 0.05 0.02 0 00:56 01:13 01:30 01:46 02:03 02:20 02:36
Active Sessions / Concurrency in-flight requests over time
in-flight (total)
12 8 4 0 00:56 01:30 02:03 02:36
System & H100 VRAM Allocation host + GPU utilization
GPU 0-3 VRAM Host CPU
100% 75% 50% 0% 00:56 01:30 02:03 02:36

โš™๏ธ Host Setup & System Hardware

Validated host CPU architecture, server type, and GPU slot enumeration.

NVIDIA_GENERIC Server Type
PASS CPU Validation
4 Slots GPU Slots
320 GB Total VRAM
๐Ÿ” Local Auth Mode Automatic Admin Bootstrap

When local authentication is selected, an administrator account is automatically created on deployment:

Username / Email: admin@yourdomain.com
Initial Password: Change_me123
Role & Privileges: SUPERADMIN / OWNER

๐Ÿงฉ GPU-Aware Model Auto-Fitting Engine

Evaluates precision quantization or lower weight variants to fit available GPU hardware.

Status Model Identifier Quant / Precision Req VRAM TP Size Recommendation / Reason
AI

Welcome to the Enterprise Data AI Portal. I have real-time access to your air-gapped Qdrant vector database and PostgreSQL data repository. How can I assist your enterprise workflow today?

TTFT: ~28ms

๐Ÿ“ฅ Ingest Enterprise Document

Chunk text and index high-dimensional embeddings into Qdrant vector storage.

๐Ÿ“š Indexed Qdrant Documents

Active vector collections in `excel_rag` Postgres & Qdrant database.

โšก vLLM Multi-GPU Tensor Parallelism Status

Qwen2.5-72B Active Model
4 GPUs Tensor Parallel Size
26.4 ms Avg TTFT
118.5 / s Tokens / Sec / Stream

๐Ÿงช Launch Multi-User Load Test

Evaluate concurrent stream processing speeds across all 4x H100 GPUs.

๐Ÿ“Š Benchmark Results

Click "Execute Async Benchmark" to measure throughput, TTFT percentiles, and TPS.