Hydra Engine
High-performance LLM inference framework for TinyLlama-1.1B with fused Triton kernels, CUDA Graphs, speculative decoding, AVX2 token verification, NF4 quantization, LibTorch, and PyBind11. Performance improvements are reported per kernel and workload rather than as one global speedup.