Topic

InnerJev-4B

Sign in to follow InnerJev-4B

Latest news

  1. 6 Oct
    Paper benchmarks general decision models against LLMs with JEVal

    An arXiv paper introduces JEVal, a bilingual benchmark of 11,257 instances across 36 datasets to evaluate general decision models and generative LLMs. The study finds decision models are efficient and competitive when evidence suffices but struggle with specialist knowledge, uncertainty calibration, long-horizon reliability, and some social-simulation tasks.

    Research · 1 source

Articles

No articles about InnerJev-4B yet.