AI Behavior LabNotebook method / EXP-GUIDE-01

Model behavior field note

Turn an interesting model response into a testable behavior question.

A research notebook method for turning questions about sycophancy, refusal consistency, bias, calibration, and framing effects into reproducible model behavior experiments.

Worked experiment frame

Sycophancy under confidence pressure

Research question
Does the candidate agree with an incorrect user claim more often when the user expresses high confidence?
Independent variable
Neutral, confident, and authority-framed versions of the same incorrect premise.
Controlled context
Same model settings, system instruction, facts, output format, and conversation length.
Observable measure
Correction, partial agreement, full agreement, uncertainty, and rationale quality.
Decision boundary
Directional finding only until repeated across topics, seeds, languages, and model versions.

Experimental practice

01

Write a falsifiable hypothesis

State what pattern you expect, which comparison could disconfirm it, and why the behavior matters. Avoid hypotheses that can be rescued after any result.

02

Change one factor at a time

Hold instructions, task facts, model settings, tools, and scoring constant while varying the pressure, framing, identity cue, or ambiguity under study.

03

Use prompt families, not one clever example

Create semantically equivalent variants across topics and phrasings. A behavior claim should survive small wording changes and should not depend on one anecdote.

04

Define labels before collecting results

Write observable scoring criteria and examples for each label. When judgment is required, preserve reviewer rationale and disagreement instead of forcing false precision.

05

Repeat and preserve configuration

Record model identifier, date, parameters, system prompt, tool state, seed where available, and every test input. Repetition helps separate a pattern from ordinary sampling variation.

06

Report limitations beside findings

Distinguish measured behavior from explanations about why it occurred. List missing variants, confounds, small samples, and contexts where the finding should not be generalized.

Preserve

Prompts, settings, labels, outputs, reviewer notes, and anomalies.

Compare

Prompt families, model versions, repeated runs, and counterexamples.

Constrain

Findings to the tested setup, sample, labels, and known limitations.