Open SourceFree on GitHub PagesGCP First

Stop Guessing Your AI Infrastructure
Start Calculating It

An open-source FinOps tool for Conversational AI pipelines. Mathematically calculates VRAM, KV Cache, end-to-end latency, and maps your STT → LLM → TTS workload to the most cost-effective cloud GPU — keeping you under the 1,500ms human-conversation threshold.

🚀 Open Calculator⭐ Star on GitHub
5
GPUs Profiled
6+
LLM Models
5
Calculator Modules
< 1,500ms
Latency Target

Five Calculators. One Tool.

Each module uses math from production deployments — not guesswork.

💾

Precision VRAM Calculator

Calculates model weight memory by quantization (FP16→INT4) with MoE active-parameter support

🧩

Smart KV Cache Estimator

Models KV cache growth with GQA & PagedAttention — reducing fragmentation from ~80% to <4%

⏱️

End-to-End Latency Simulator

Breaks down STT-LLM-TTS pipeline against the 1,500ms human-conversation threshold

🖥️

Hardware Matchmaker

Auto-selects the most cost-efficient GPU (T4 → L4 → A100 → H100) for your VRAM & CCU target

💰

FinOps Blueprint

Compares on-demand vs spot pricing and generates a Kubernetes HPA config based on queue depth

📊

CCU Capacity Planner

Estimates concurrent capacity units per GPU for STT (RTF-based) and LLM (TPS-based)

🤖 Supported LLM Models

Llama 3 8BLlama 3 70BGemma 2 9BQwen 2.5 7BMixtral 8x7BMistral 7B+ more

🖥️ Profiled GPUs (GCP/AWS/Azure)

NVIDIA T4NVIDIA L4A100 40GBA100 80GBH100 80GB
🎯

Ready to size your pipeline?

Enter your model, target CCU, and latency constraints. Get GPU recommendations with monthly cost estimates in seconds.

Launch Free Calculator →