VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO

(arxiv.org)

344 points | by timhigins 17 hours ago ago

113 comments