LLM fine-tuning
Mini-LLM Puzzle Solver
A 3B-parameter LLaMA fine-tuned on NYT Connections puzzles with hybrid SFT + GRPO, improving perfect-solve rate and cluster purity.
More than doubledperfect-solve rate; cluster purity improved
Approach
- Fine-tuned a 3B-parameter LLaMA on NYT Connections puzzles.
- Hybrid SFT + GRPO training.
Results
- More than doubled perfect-solve rate; improved cluster purity.
Stack
- PyTorch
- Hugging Face