Field Guide

Define what good AI work looks like before you delegate

Research-based guidance; hypothetical example

Before asking an AI assistant to prepare a piece of work, write down what you will inspect when it returns. For a meeting summary, that might be whether each action has a named owner or an explicit “owner not assigned.” For a comparison, it might be whether the recommendation changes when a required feature is missing.

These are acceptance criteria: conditions the result must meet for you to use it. A useful criterion names the requirement, the evidence, and who will judge it. “Make this useful” leaves all three undecided.

This guide gives you a worksheet and a worked example for defining those conditions before delegation. It is research-based guidance with an original hypothetical example, not a method Today’s Worker has validated through repeated use.

Choose the decision the output must support

Start with the person receiving the work and the next decision they need to make. “Summarize these notes” could mean a record of the discussion, a list of assignments, or a recommendation for the next meeting. Each needs different checks.

Try this sentence: “The reader needs to use this output to ___.” Then name the deliverable. In our hypothetical example, a small team has supplied notes from a project meeting. Its coordinator needs an action list to prepare the next check-in. The assistant will produce a table of agreed actions and a separate list of unresolved questions.

The acceptance decision is whether that handoff is faithful and usable. Whether the team later completes its actions is a separate outcome. A good-looking table cannot establish that the project is progressing.

Pair each requirement with a way to check it

For every requirement, finish the sentence “I will verify this by ___.” If you cannot finish it, narrow the requirement or acknowledge that it needs judgment.

Some checks are straightforward: each row contains an action, owner, due date, and source reference. Other checks require reading: an action preserves the agreement in the notes without making it stronger. A complete row can still misrepresent the meeting.

Anthropic’s guide to agent evaluations distinguishes the assistant’s account of its work from the outcome being evaluated. It also recommends clear tasks whose grading requirements are visible in the instructions. Applied here, the check is inspecting the delivered action list against the notes; the assistant’s statement that it checked everything is insufficient.

Keep essential requirements separate from preferences. A missing owner must be visible. Whether the table uses sentence case is a preference unless your receiving system requires it. Avoid letting an attractive format compensate for an inaccurate assignment.

Work through an example before sending the task

Here are three invented meeting notes. They contain no real team or client information:

A suitable action list would contain the revised brief assigned to Maya, due “Tuesday; calendar date not supplied,” with reference N1. It would contain the signup-link repair with “unassigned” and “before launch; date unresolved,” with reference N3. Customer interviews belong under unresolved questions because N2 records a suggestion, not a commitment.

Now write the criteria that distinguish that result from a plausible but wrong one:

  1. Commitments: Include both agreed actions, N1 and N3. Keep the unapproved interview idea out of the action table. Check each input note against its treatment in the output.
  2. Attribution: Assign Maya only to the revised brief. Mark the signup repair unassigned. Compare each owner field with its supporting note.
  3. Dates: Preserve the supplied timing without inventing calendar dates. Check that Tuesday and the unresolved launch dependency remain visible.
  4. Traceability: Attach the supporting note ID to every action. Open that note and verify that it actually supports the row.
  5. Usefulness: The coordinator can tell what needs follow-up without rereading all the notes. The coordinator judges this by identifying the missing repair owner, launch date, and decision on interviews from the output.

These checks do not prescribe a particular sentence or row order. Several outputs could pass. The example establishes the meaning that must survive the transformation.

Specify what an incomplete answer should look like

Missing information can make an honest output less tidy. Decide how it should appear before the assistant fills the gap.

For the meeting example, unknown owners and dates do not prevent a useful draft. The assistant should label them and include them in the follow-up list. If the notes themselves are absent or unreadable, it cannot produce a faithful action list; it should identify the missing input.

A conflict also needs visible treatment. If another supplied note assigns the same repair to a different person, the output should identify the conflict and cite both notes. Your criterion should say whether a newer approved record resolves that conflict. Without that rule or evidence, choosing an owner would hide the uncertainty.

Copy this acceptance worksheet

Fill this in for one deliverable. Use as many requirement lines as the task needs, but make each earn its place.

Reader and next decision: Who will use the result, and for what?

Deliverable: What artifact should be returned?

Inputs: Which supplied materials define the facts? How will they be identified?

Required condition: What must be true of the result?

Verification: What will I inspect, compare, or calculate to check that condition?

Judgment: What needs a person’s assessment, and who will make it?

Missing or conflicting information: What should be labeled, omitted, or returned as unresolved?

Acceptance record: For each condition, record pass, revise, or unable to verify, with a short reason.

Attach the completed worksheet to the task request. For the example above, add: “Produce the action table and unresolved questions using notes N1–N3 and these criteria. Include the note references with the result.” The criteria belong in the request the assistant actually receives.

Review the result without moving the target

Inspect the returned artifact using the checks you wrote. If a date is invented, point to the failed date criterion and the supporting note. If you discover that you also need action priorities, record that as a new requirement. This makes the difference between a missed instruction and a changed request visible.

OpenAI’s evaluation guidance recommends evaluations specific to the task and human feedback to calibrate automated scoring. For everyday work, you can begin with a manual review of the worksheet. An AI reviewer may help locate possible defects, but its judgment still needs checking against the actual material.

Keep the scope modest. One accepted action list establishes that this result met these checks. It does not establish dependable performance on all meeting notes. Nor will a checklist rescue incomplete source material or a decision that requires expertise the reviewer lacks.

Take one task you plan to delegate today and write its acceptance worksheet before running it. Save the returned artifact and your review beside the request. If the same failure returns across tasks, use the guide to turn recurring AI mistakes into reusable instructions. The worksheet defines what this task needs; the later correction helps you decide what should carry forward.

Sources and further reading