0wr asked which harbor A/Bs were decided by a narrow margin before 39.4. Measuring harbor's variance makes the margin-by-margin triage moot. Harbor's paired seed-to-seed sd is 6.19 fails (24 paired ON/OFF runs, budget 2500), so n=3 resolves nothing finer than 15.4 fails. Every recorded harbor margin is below that: 13.9 share_edge_cap 3.7, 20 qpk 8.3, 23 f1d mixed, 37.1 tiering 6.3. So it is not that SOME harbor results were narrow -- no harbor A/B run at three seeds could resolve the margin it reported, independently of what 39.4 did to the programme. Of the 220 possible 3-seed subsets of the 24 runs, 56 (25%) show a clean 3/3 sweep for ON. Re-measured 20's harbor arm, the one backing a live default: N=3 2W/1L/0T +2.67 p=0.560 N=12 8W/3L/1T +3.50 p=0.076 N=24 13W/10L/1T +1.21 p=0.502 CI [-2.46,+4.88] Null. The published "harbor: ON wins 3/3, 80.3 -> 72.0" was a lucky draw -- even seeds 1-3 measured here give 2W/1L, not a sweep. So 20's claim that the qpk verdict "holds at both example scales tested" is withdrawn and annotated in place. collapse_insearch's default rests on programme-house alone (38.19, N=60, +0.57, p=0.017). It is not refuted on harbor -- direction positive but indistinguishable from zero -- but harbor must not be cited as corroboration. Harness generalised (PROG/BUDGET/WORKERS) and results kept. Filed homemaker-py-... : A/B harnesses should report the minimum detectable difference for the N they run, so an underpowered verdict is visible when it is made rather than years later. Closes homemaker-py-0wr. Lint at parity (46). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB |
||
|---|---|---|
| .. | ||
| hooks | ||
| .gitignore | ||
| config.yaml | ||
| issues.jsonl | ||
| metadata.json | ||
| README.md | ||
Beads - AI-Native Issue Tracking
Welcome to Beads! This repository uses Beads for issue tracking - a modern, AI-native tool designed to live directly in your codebase alongside your code.
What is Beads?
Beads is issue tracking that lives in your repo, making it perfect for AI coding agents and developers who want their issues close to their code. No web UI required - everything works through the CLI and integrates seamlessly with git.
Learn more: github.com/steveyegge/beads
Quick Start
Essential Commands
# Create new issues
bd create "Add user authentication"
# View all issues
bd list
# View issue details
bd show <issue-id>
# Update issue status
bd update <issue-id> --claim
bd update <issue-id> --status done
# Sync with Dolt remote
bd dolt push
Working with Issues
Issues in Beads are:
- Git-native: Stored in Dolt database with version control and branching
- AI-friendly: CLI-first design works perfectly with AI coding agents
- Branch-aware: Issues can follow your branch workflow
- Always in sync: Auto-syncs with your commits
Why Beads?
✨ AI-Native Design
- Built specifically for AI-assisted development workflows
- CLI-first interface works seamlessly with AI coding agents
- No context switching to web UIs
🚀 Developer Focused
- Issues live in your repo, right next to your code
- Works offline, syncs when you push
- Fast, lightweight, and stays out of your way
🔧 Git Integration
- Automatic sync with git commits
- Branch-aware issue tracking
- Dolt-native three-way merge resolution
Get Started with Beads
Try Beads in your own projects:
# Install Beads
curl -sSL https://raw.githubusercontent.com/steveyegge/beads/main/scripts/install.sh | bash
# Initialize in your repo
bd init
# Create your first issue
bd create "Try out Beads"
Learn More
- Documentation: github.com/steveyegge/beads/docs
- Quick Start Guide: Run
bd quickstart - Examples: github.com/steveyegge/beads/examples
Beads: Issue tracking that moves at the speed of thought ⚡