NInfer Multivariate Optimization Studio RTX 5090 • NVFP4
Empirical 45-run multi-scenario sweep & DFlash-2 speculative telemetry analysis
Engine: NInfer (CUDA 13.1 / SM 12.0)
Model: Qwen 3.8 27B NVFP4 (Swift 1.5)
Champion: Config-1.4 (Tight Agentic Budget)
DFlash-2 Peak Decode
350+ tok/s
7 draft tokens + LM head draft
Speculative Acceptance
62.6%
Across 3,549 real-world DFlash-2 requests
Prefix Cache Hit Rate
96.4%
Avg TTFT 0.15s – 0.35s under 60k context
Optimal Agentic Profile
T=0.65, B=1200
98% tail-latency reduction (Config-1.4)
Parallel Coordinates: 6D Hyperparameter Flow to Optimization Loss
Traces how Architecture, Temperature, Presence Penalty, Thinking Budget, and Latency map directly to Optimization Loss.
🏆 Cyan Bundle = Optimal Paths
Loss Landscape Contour (Temperature vs. Thinking Budget)
Visualizes the "Optimal Valley" around (T=0.65, B=1200) vs. high-penalty runaway cliffs.
Multi-Scenario Model Radar Signatures
Comparing DFlash-2 vs. Swift10-MTP vs. Swift15-MTP across bug fixing, hallucination traps, and synthesis.
DFlash-2 Draft Position Acceptance Decay (Pos 1 – 7)
Acceptance rate per speculative draft token index across 3,549 completed DFlash-2 requests.
DFlash-2 Acceptance Rate vs. Context Length
Shows speculative efficiency scaling from small prompts up to 64,000 token context.
Time to First Token (TTFT) vs. Prompt Tokens (Prefix Cache Scaling)
Real-world TTFT response curve demonstrating response-replay cache acceleration.
Decode Throughput Distribution (Tokens / Second)
Empirical token generation speed distribution on RTX 5090 Blackwell NVFP4.
Pareto Efficiency Frontier: Quality Score vs. Total Wallclock Time
Bubble size represents Thinking Tokens. Cyan dashed line represents the optimal non-dominated boundary.
Cumulative 45-Run Agentic Benchmark Scorecard
Filterable and searchable results table across all configurations and scenarios.
Scenario Config Family Sampling (T / P / Pen / Budg) Quality Opt Loss Wallclock TTFT Thinking Loops