b1057f6105
fitglme coefTest only offers observation-level DF (df2~=178), which overstates the group x day interaction significance for this few-subject design (p~=0 in every scenario). Replace the reported learning-rate result with a cluster-honest per-animal slope test (per-subject OLS slope of success vs day, Welch + MWU, B2 vs A2) -- now correctly parallel in all 6 scenarios, matching the Python GEE. The GLMM F-test is kept as a flagged reference only. Adds a regression test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>