-
MMLU-Pro Leaderboard
π₯254More advanced and challenging multi-task evaluation
-
Stick To Your Role! Leaderboard
π63Benchmarking LLMs on the stability of simulated populations
-
ZeroEval Leaderboard
π53Explore ZeroEval embedding benchmark online
-
Open Medical-LLM Leaderboard
π₯437Explore and submit models for benchmarking
Hristo Panev
hppdqdq
AI & ML interests
None yet
Recent Activity
liked a model about 3 hours ago
MiniMaxAI/MiniMax-Music3 liked a model 8 days ago
ReadyArt/gemma-4-31B-it-scotoma-GGUF liked a model about 2 months ago
tsolful/Krea2_Turbo_Raw_INT8Organizations
None yet