bitnet triton
Custom Triton kernel for 1.58-bit LLM inference โ 4.4ร less VRAM, 1.5ร faster decode, +0.26% PPL on RTX 4060 Laptop
Python
Custom Triton kernel for 1.58-bit LLM inference โ 4.4ร less VRAM, 1.5ร faster decode, +0.26% PPL on RTX 4060 Laptop
Cross-platform Material-Design chat app with a threaded Python socket server, SQLite persistence, and PBKDF2-hashed accoโฆ
Custom CNN + Bidirectional LSTM speech recognition model in TensorFlow, packaged for distributed training on Google Clouโฆ