Estimate model-weight and KV-cache footprint per model, and see how many concurrent users fit before the card runs out, for Ollama, Llama CPP or vLLM deployments sharing one GPU.
vLLM: set --gpu-memory-utilization per instance so the sum across instances stays under 1.0 minus overhead.
GPU allocation summary
62.1 / 96 GB used
Model weightsKV cacheOverhead reserveFree
Models on this GPU
Presets are approximate architecture values — verify against the model's config.json.
Weights17.6 GB
KV / user2.10 GB
KV total (15 users)31.5 GB
Model total49.1 GB
Decode (1 user)102 tok/s
Aggregate (15 users)1527 tok/s
Time to first token1016 ms
Compute ceiling7875 tok/s
Presets are approximate architecture values — verify against the model's config.json.
Weights1.2 GB
KV / user0.02 GB
KV total (15 users)0.4 GB
Model total1.6 GB
Decode (1 user)1493 tok/s
Aggregate (15 users)22400 tok/s
Time to first token5 ms
Compute ceiling420000 tok/s
Client Success Stories
Transforming Businesses Through Intelligence
See how forward-thinking companies leverage autonomous systems and AI-driven solutions to revolutionize their operations
RW
Rebecca Walsh
Head of Digital Transformation
Dennis delivered an autonomous agent ecosystem that revolutionized our customer service operations. The AI systems learn and adapt in real-time, providing personalized responses that our customers love. We've seen a 300% improvement in customer satisfaction scores.
AI Customer Experience Platform
JM
James Mitchell
CTO, Financial Services
The intelligent workflow automation Dennis built for us processes millions of transactions daily with zero human intervention. His systems don't just follow rules—they understand context and make decisions like our best analysts would.
Autonomous Trading Systems
LC
Lisa Chen
VP Engineering, HealthTech
Working with Dennis transformed our entire DevOps culture. His AI-powered deployment pipelines predict and prevent issues before they occur. We've achieved 99.99% uptime while deploying 10x more frequently than before.