$cd /research|AI Infrastructure
Research /systems
$ Investigating AI infrastructure, LLM optimization, and intelligent systems. Publications, projects, and technical deep dives.
Projects
─4 active01
GGUF Memory Optimization
Ongoing2026A systematic approach to estimating and optimizing VRAM usage for GGUF models in production environments.
#GGUF#VRAM#Optimization#LLM
02
Distributed Inference Systems
In Review2026Design patterns for scalable LLM inference across heterogeneous GPU clusters with minimal latency.
#Distributed Systems#Inference#Scalability
03
KV Cache Quantization Strategies
Published2026Comparative analysis of quantization techniques for KV cache in transformer models and their impact on memory usage.
#KV Cache#Quantization#Memory#Transformers
04
AI Agent Orchestration Frameworks
Ongoing2026Evaluating orchestration patterns for multi-agent AI systems in enterprise workflows and decision-making.
#AI Agents#Orchestration#Workflows#Enterprise
Publications
─3 papersEfficient Memory Management for Large Language Models
International Conference on Machine Learning Systems (MLSys)|Conference Paper|2025
Read
Grouped-Query Attention: Optimization Techniques
NeurIPS Workshop on Efficient AI|Workshop Paper|2024
Read
Scaling AI Inference: A Practical Guide
arXiv Preprint|Technical Report|2025
Read
Research Areas
AI InfrastructureLLM OptimizationGPU Memory ManagementDistributed SystemsAI AgentsPerformance EngineeringSystem DesignMachine Learning
$ Collaborating on AI research?
Get in touch