EVO CAPABILITIES / Self-checking and correction loops

Make a correction useful beyond this reply.

EVO connects outcome checks with reviewable improvement suggestions, so a correction can shape future assistance without silently changing the product or your permissions.

Mac development · See current scope

The real problem

When an assistant misses your request, fixing the immediate answer is only part of the job. You may also want it to stop repeating the same presentation mistake or to check that a requested result actually happened. Repeating a preference in every chat is frustrating; silently turning one complaint into a permanent rule can be just as wrong.

What chat history leaves you to do

A conversational apology can acknowledge a mistake without changing anything that affects the next task. A blanket instruction to learn from every exchange has the opposite problem: it can capture a temporary requirement, misunderstand the correction or treat the user's frustration as a diagnosis of the cause. Useful improvement needs a visible path from a real correction to a specific change.

How EVO approaches it

  1. Check outcomes, not just confident wording

    EVO's work review separates available results from missing checks. A tool's error cannot stand in for evidence from a successful result, and a completed response cannot accept the whole goal by itself. These checks are intended to make incomplete work easier to identify and repair before an answer presents it as settled.

  2. Capture a reported mismatch carefully

    The local improvement workflow can recognise supported direct corrections after a saved conversation turn. It links a concise candidate summary to the original work instead of copying the full conversation into another log. The candidate records that something may have missed the mark; it does not claim to know the true cause or decide who is to blame.

  3. Turn repeated needs into a reviewable suggestion

    You can inspect the source, confirm or correct the category, and review a supported adjustment. Examples include leading with the result, reducing repeated background or checking the requested scope before answering. These are specific preferences and final checks. A change of requirements can remain a note rather than become a new permanent behaviour.

  4. Keep improvement reversible

    Only an explicitly approved suggestion becomes active guidance for later supported responses. You can revoke it, delete its source record or turn capture off. Deleting the improvement history disables further capture until re-enabled. Product-wide code changes, new tools and model training are different activities and cannot be authorised by a correction record.

A concrete example

ILLUSTRATIVE WORKFLOW · NOT A CUSTOMER RESULT

Illustrative workflow: stop burying the answer

You repeatedly tell EVO that you need the decision first and the supporting explanation afterward. The original task still needs a corrected answer now.

  1. Correct the current response while preserving the original request and its important details.

  2. Review the local candidate linked to that correction and confirm that it is about presentation.

  3. If a supported result-first suggestion is proposed, approve it for later work and revoke it if it no longer helps.

What you take away

The intended benefit is a preference you can see and control, rather than an unexplained promise that the model has learned. Later responses can receive the approved guidance while detailed requests and necessary verification still take priority.

Your choices

  • Inspect the source of a correction, correct its category and decide whether a proposed adjustment is useful.
  • Approve a specific suggestion, revoke active guidance or pause local capture.
  • Delete improvement records without pretending the original conversations or their separate backups have also been erased.

Current scope

Implemented foundation

  • Controlled tests preserve completed operation results when a later response fails and prevent automatic replay of uncertain actions.
  • Bounded recovery is not general self-improvement. Cognitive research stays off or Shadow; it does not train the model or change a public user’s permissions.

Next milestones

  • Public availability and real-user evidence of sustained improvement remain unfinished. Recognising a correction does not prove its root cause or guarantee that the next model response follows the guidance.
  • EVO does not retrain model weights, install new skills, rewrite its own code or deploy product changes through this feature. Broader improvements require separate engineering review and explicit authorisation.

These are development capabilities, not a public release. See platform availability before requesting a trial.

Platforms and availability

Related questions

What if a tool succeeds but my goal is not achieved?

A successful tool result should count as evidence of that operation, not proof that the whole task is complete.

EVO’s supported workflows distinguish a proposed change, an applied change and an outcome that has been checked against the original request.

Saving a new button label proves the file changed; it does not prove visitors understand it or that the button works on mobile.

Inspect the saved result, identify the unmet condition and request the next focused check or correction.

The system can preserve this distinction, but model judgment may still miss a requirement. Important outcomes need direct verification.

What happens if EVO guesses my intention incorrectly?

Your explicit correction should take priority over the conflicting interpretation while preserving requirements you have not changed.

EVO can revise the brief and produce a more appropriate next result. A tentative interpretation should remain separate from a confirmed user fact.

If “simplify” meant fewer steps rather than shorter text, say so; the next result should address the workflow instead of repeatedly trimming words.

State the intended outcome or answer the relevant clarification. You can also review a captured correction separately.

A correction does not guarantee the model will reason perfectly next time, and it cannot undo an already completed external action.

How can EVO learn from my corrections?

The local improvement feature can turn relevant direct corrections into reviewable suggestions for future response style and final checks.

It records a cautious summary of the reported mismatch. An adjustment becomes active only after you confirm the relevant records and approve the proposed change.

Repeated requests to lead with the result can become a reviewed preference for placing the useful answer before background.

Inspect, correct or delete the records; approve or revoke supported adjustments and disable capture.

This changes selected response guidance, not model weights, product code or tool permissions. Long-term improvement in real-user outcomes has not yet been demonstrated.

What does EVO check before presenting a result?

Supported workflows check the result against its required structure, available evidence and requested outcome, keeping unverified parts separate from observed results.

Research, selected-file changes and tool results use different checks. Additional reasoning can explain a problem, but cannot substitute for a missing test or source.

A proposed website file can pass source checks while the final report still says its rendered layout has not been reviewed.

Open the checks, ask for a focused follow-up and decide whether further verification is necessary.

Self-checking reduces opportunities for unsupported claims but is not proof that every answer is correct. Model review and user acceptance have different roles.

Try the workflow

Bring a task you want to move forward.

Explore the Mac release, or tell us which workflow you would like to evaluate. An application does not guarantee an invitation.

Get EVO AlphaShare your workflow