Topic
InnerJev-4B
Latest news
- 6 OctPaper benchmarks general decision models against LLMs with JEVal
An arXiv paper introduces JEVal, a bilingual benchmark of 11,257 instances across 36 datasets to evaluate general decision models and generative LLMs. The study finds decision models are efficient and competitive when evidence suffices but struggle with specialist knowledge, uncertainty calibration, long-horizon reliability, and some social-simulation tasks.
Research · 1 source
Articles
No articles about InnerJev-4B yet.