-
Discovering Language Model Behaviors with Model-Written Evaluations
Authors:
Ethan Perez,
Sam Ringer,
Kamilė Lukošiūtė,
Karina Nguyen,
Edwin Chen,
Scott Heiner,
Craig Pettit,
Catherine Olsson,
Sandipan Kundu,
Saurav Kadavath,
Andy Jones,
Anna Chen,
Ben Mann,
Brian Israel,
Bryan Seethor,
Cameron McKinnon,
Christopher Olah,
Da Yan,
Daniela Amodei,
Dario Amodei,
Dawn Drain,
Dustin Li,
Eli Tran-Johnson,
Guro Khundadze,
Jackson Kernion
, et al. (38 additional authors not shown)
Abstract:
As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evaluations with crowdwork (which is time-consuming and expensive) or existing data sources (which are not always available). Here, we automatically generate evaluations with LMs. We explore approaches with varying amounts of human effort, from inst…
▽ More
As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evaluations with crowdwork (which is time-consuming and expensive) or existing data sources (which are not always available). Here, we automatically generate evaluations with LMs. We explore approaches with varying amounts of human effort, from instructing LMs to write yes/no questions to making complex Winogender schemas with multiple stages of LM-based generation and filtering. Crowdworkers rate the examples as highly relevant and agree with 90-100% of labels, sometimes more so than corresponding human-written datasets. We generate 154 datasets and discover new cases of inverse scaling where LMs get worse with size. Larger LMs repeat back a dialog user's preferred answer ("sycophancy") and express greater desire to pursue concerning goals like resource acquisition and goal preservation. We also find some of the first examples of inverse scaling in RL from Human Feedback (RLHF), where more RLHF makes LMs worse. For example, RLHF makes LMs express stronger political views (on gun rights and immigration) and a greater desire to avoid shut down. Overall, LM-written evaluations are high-quality and let us quickly discover many novel LM behaviors.
△ Less
Submitted 19 December, 2022;
originally announced December 2022.
-
Deepchecks: A Library for Testing and Validating Machine Learning Models and Data
Authors:
Shir Chorev,
Philip Tannor,
Dan Ben Israel,
Noam Bressler,
Itay Gabbay,
Nir Hutnik,
Jonatan Liberman,
Matan Perlmutter,
Yurii Romanyshyn,
Lior Rokach
Abstract:
This paper presents Deepchecks, a Python library for comprehensively validating machine learning models and data. Our goal is to provide an easy-to-use library comprising of many checks related to various types of issues, such as model predictive performance, data integrity, data distribution mismatches, and more. The package is distributed under the GNU Affero General Public License (AGPL) and re…
▽ More
This paper presents Deepchecks, a Python library for comprehensively validating machine learning models and data. Our goal is to provide an easy-to-use library comprising of many checks related to various types of issues, such as model predictive performance, data integrity, data distribution mismatches, and more. The package is distributed under the GNU Affero General Public License (AGPL) and relies on core libraries from the scientific Python ecosystem: scikit-learn, PyTorch, NumPy, pandas, and SciPy. Source code, documentation, examples, and an extensive user guide can be found at \url{https://github.com/deepchecks/deepchecks} and \url{https://docs.deepchecks.com/}.
△ Less
Submitted 16 March, 2022;
originally announced March 2022.
-
Quantum circuits with many photons on a programmable nanophotonic chip
Authors:
J. M. Arrazola,
V. Bergholm,
K. Brádler,
T. R. Bromley,
M. J. Collins,
I. Dhand,
A. Fumagalli,
T. Gerrits,
A. Goussev,
L. G. Helt,
J. Hundal,
T. Isacsson,
R. B. Israel,
J. Izaac,
S. Jahangiri,
R. Janik,
N. Killoran,
S. P. Kumar,
J. Lavoie,
A. E. Lita,
D. H. Mahler,
M. Menotti,
B. Morrison,
S. W. Nam,
L. Neuhaus
, et al. (14 additional authors not shown)
Abstract:
Growing interest in quantum computing for practical applications has led to a surge in the availability of programmable machines for executing quantum algorithms. Present day photonic quantum computers have been limited either to non-deterministic operation, low photon numbers and rates, or fixed random gate sequences. Here we introduce a full-stack hardware-software system for executing many-phot…
▽ More
Growing interest in quantum computing for practical applications has led to a surge in the availability of programmable machines for executing quantum algorithms. Present day photonic quantum computers have been limited either to non-deterministic operation, low photon numbers and rates, or fixed random gate sequences. Here we introduce a full-stack hardware-software system for executing many-photon quantum circuits using integrated nanophotonics: a programmable chip, operating at room temperature and interfaced with a fully automated control system. It enables remote users to execute quantum algorithms requiring up to eight modes of strongly squeezed vacuum initialized as two-mode squeezed states in single temporal modes, a fully general and programmable four-mode interferometer, and genuine photon number-resolving readout on all outputs. Multi-photon detection events with photon numbers and rates exceeding any previous quantum optical demonstration on a programmable device are made possible by strong squeezing and high sampling rates. We verify the non-classicality of the device output, and use the platform to carry out proof-of-principle demonstrations of three quantum algorithms: Gaussian boson sampling, molecular vibronic spectra, and graph similarity.
△ Less
Submitted 2 March, 2021;
originally announced March 2021.
-
Comment on "Non-representative Quantum Mechanical Weak Values"
Authors:
Alon Ben Israel,
L. Vaidman
Abstract:
Svensson [Found. Phys. 45, 1645 (2015)] argued that the concept of the weak value of an observable of a pre- and post-selected quantum system cannot be applied when the expectation value of the observable in the initial state vanishes. Svensson's argument is analyzed and shown to be inconsistent using several examples.
Svensson [Found. Phys. 45, 1645 (2015)] argued that the concept of the weak value of an observable of a pre- and post-selected quantum system cannot be applied when the expectation value of the observable in the initial state vanishes. Svensson's argument is analyzed and shown to be inconsistent using several examples.
△ Less
Submitted 25 August, 2016; v1 submitted 25 August, 2016;
originally announced August 2016.