What works, what goes wrong, and how to tell the difference.
The one place delegation is unambiguous. There is a red test, there is a green test, and the agent can tell which it has. Flint reads the failure, edits, and runs it again until it passes or it explains why it cannot.
The rule that matters more than the tool: give the agent something that can prove it succeeded. A task with a test, a build or a linter attached comes back verified. A task with none comes back plausible, and plausible is the expensive kind of wrong.
A desktop AI agent for Windows, built in Auckland, New Zealand. It reads your files, runs your terminal and drives your browser. Billed per token with a free monthly allowance and no subscription. See pricing, the FAQ, or who builds it.