An open-source FinOps tool for Conversational AI pipelines. Mathematically calculates VRAM, KV Cache, end-to-end latency, and maps your STT → LLM → TTS workload to the most cost-effective cloud GPU — keeping you under the 1,500ms human-conversation threshold.
Each module uses math from production deployments — not guesswork.
Calculates model weight memory by quantization (FP16→INT4) with MoE active-parameter support
Models KV cache growth with GQA & PagedAttention — reducing fragmentation from ~80% to <4%
Breaks down STT-LLM-TTS pipeline against the 1,500ms human-conversation threshold
Auto-selects the most cost-efficient GPU (T4 → L4 → A100 → H100) for your VRAM & CCU target
Compares on-demand vs spot pricing and generates a Kubernetes HPA config based on queue depth
Estimates concurrent capacity units per GPU for STT (RTF-based) and LLM (TPS-based)
Enter your model, target CCU, and latency constraints. Get GPU recommendations with monthly cost estimates in seconds.
Launch Free Calculator →