CogBench: a large language model walks into a psychology lab

Coda-Forno, Julian; Binz, Marcel; Wang, Jane X.; Schulz, Eric

Computer Science > Computation and Language

arXiv:2402.18225 (cs)

[Submitted on 28 Feb 2024]

Title:CogBench: a large language model walks into a psychology lab

Authors:Julian Coda-Forno, Marcel Binz, Jane X. Wang, Eric Schulz

View PDF HTML (experimental)

Abstract:Large language models (LLMs) have significantly advanced the field of artificial intelligence. Yet, evaluating them comprehensively remains challenging. We argue that this is partly due to the predominant focus on performance metrics in most benchmarks. This paper introduces CogBench, a benchmark that includes ten behavioral metrics derived from seven cognitive psychology experiments. This novel approach offers a toolkit for phenoty** LLMs' behavior. We apply CogBench to 35 LLMs, yielding a rich and diverse dataset. We analyze this data using statistical multilevel modeling techniques, accounting for the nested dependencies among fine-tuned versions of specific LLMs. Our study highlights the crucial role of model size and reinforcement learning from human feedback (RLHF) in improving performance and aligning with human behavior. Interestingly, we find that open-source models are less risk-prone than proprietary models and that fine-tuning on code does not necessarily enhance LLMs' behavior. Finally, we explore the effects of prompt-engineering techniques. We discover that chain-of-thought prompting improves probabilistic reasoning, while take-a-step-back prompting fosters model-based behaviors.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2402.18225 [cs.CL]
	(or arXiv:2402.18225v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2402.18225

Submission history

From: Julian Coda-Forno [view email]
[v1] Wed, 28 Feb 2024 10:43:54 UTC (755 KB)

Computer Science > Computation and Language

Title:CogBench: a large language model walks into a psychology lab

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:CogBench: a large language model walks into a psychology lab

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators