Topic

JEVal

Sign in to follow JEVal

Latest news

  1. 6 Oct
    Paper benchmarks general decision models against LLMs with JEVal

    An arXiv paper introduces JEVal, a bilingual benchmark of 11,257 instances across 36 datasets to evaluate general decision models and generative LLMs. The study finds decision models are efficient and competitive when evidence suffices but struggle with specialist knowledge, uncertainty calibration, long-horizon reliability, and some social-simulation tasks.

    Research · 1 source

Articles

No articles about JEVal yet.