03 / WRITING
Notes, ideas, and learnings
Years of thinking out loud — from reinforcement learning and graph learning, through T5 and SSMs, to modern LLM research and systems thinking. The archive is kept whole on purpose.
2025
- RL Bite: Monte Carlo Search Tree
- RL Bite: Monotonic Policy Improvement and Deriving Proximal Policy Optimization (PPO)
- RL Bite: Policy Gradient and Reinforce
- RL Bite: Learning the Q Function
- TLDR; Graph Contrastive Learning: Representation Scattering
- RL Bite: Computing the Value Function
- TLDR; HC-GAE The Hierarchical Cluster-based Graph Auto-Encoder for Graph Representation Learning
- RL Bite: Bellmans Equations and Value Functions
- TLDR; Duplex: Dual GAT for Complex Embeddings of Directed Graphs
- RL Bite: Exploitation vs Exploration
2024
- Graph Neural Networks meet Large Language Models
- Hymba, a new breed of SSM-Attention Hybrids
- Transform any LLMs to a powerful Encoder
- 2025 Year of Zig
- Distilling State Space Models from Transformers
- Illusion of State in SSMs like Mamba
- Mamba(2) and Transformer Hybrids: An Overview
- Hydra a Double Headed Mamba
- From Mamba to Mamba-2
- Butterflies, Monarchs, Hyenas, and Lightning Fast BERT
- BinT5 and HexT5 or T5 and Binary Reverse Engineering
- CodeT5 and CodeT5+
- Longer Context for T5
- T5 the Old New Thing