LLM fine-tuning

Mini-LLM Puzzle Solver

A 3B-parameter LLaMA fine-tuned on NYT Connections puzzles with hybrid SFT + GRPO, improving perfect-solve rate and cluster purity.

More than doubledperfect-solve rate; cluster purity improved

Approach

  • Fine-tuned a 3B-parameter LLaMA on NYT Connections puzzles.
  • Hybrid SFT + GRPO training.

Results

  • More than doubled perfect-solve rate; improved cluster purity.

Stack

  • PyTorch
  • Hugging Face