LocoMusa
Not "how smart is the model?" but "how good is it to think with?" A small, curated benchmark for local LLMs scoring ideation, reframing, steelmanning, and multi-turn dialogue — by blind pairwise judging against frontier-model anchors.
Not "how smart is the model?" but "how good is it to think with?" A small, curated benchmark for local LLMs scoring ideation, reframing, steelmanning, and multi-turn dialogue — by blind pairwise judging against frontier-model anchors.