AffixIO Skill Lab / agent workflow evaluation

Teach agents to finish the workflow, then prove what happened.

SkillGym turns human-written skills into executable tasks, outcome checkers and contrastive runs. AffixIO applies the part that matters at the action boundary: make the workflow explicit, verify the result and preserve evidence before an agent moves on.

Task instanceCode-style checkerBaseline vs skillAffixIO evidence

From a skill card to a decision you can inspect.

A prompt can describe a process. It cannot, by itself, show that the process completed. Skill Lab separates the reusable workflow from the concrete task, runs a checker against the requested outcome and records the boundary conditions that produced the result.

01 / CARD

Write the workflow

Describe the tools, order of operations, constraints and stop conditions an agent must follow.

02 / TASK

Instantiate a case

Turn that reusable skill into a concrete merchant, tool call or eligibility scenario.

03 / CHECK

Verify the outcome

Check the final state, schema, limits and evidence instead of grading the agent’s prose.

04 / RECORD

Keep the boundary

Bind the action, decision and digest so another system can inspect what was allowed.

Inspired by SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving. AffixIO is not affiliated with the authors or the SkillGym project.

Run a contrastive skill check.

Choose a workflow, then compare a bare baseline with a guided run. The checker rewards the conditions that matter, not a particular answer style. Everything below runs in your browser and uses synthetic data.

Ready for a taskIdle

Choose a task to create a synthetic environment.

    AffixIO evidence record / synthetic demo

    Where this meets AffixIO.

    SkillGym focuses on turning procedural experience into training data. AffixIO focuses on the verifiable boundary around a real action. Together, the ideas point to a practical agent architecture.

    Skills are not authority

    A good workflow can teach an agent how to act. It should not silently expand what the agent is allowed to do. Keep merchant, amount, expiry and approval scope separate.

    Checkers beat confident text

    Use executable checks for the final state, input schema and policy constraints. Then return a clear allow, review or deny result to the next system.

    Trajectories need context

    A useful record includes the task, action, constraints, decision and digest. AffixIO can sign the outcome and connect it to an audit path.

    This page does not fine-tune a model, collect training trajectories or grant production payment authority. It demonstrates a design pattern: procedural skill plus executable verification plus a bounded, inspectable decision.

    Build an agent that can show its work.

    Explore the AffixIO SDK, agent verification and Agentic Pay Kit for the production pieces behind this lab.

    Explore Agentic Pay Kit

    Free proof allocations

    Agents get 150 free proofs. Humans get 100. No card required.

    AI agents receive 150 free AffixIO SDK proofs on BoundProof Agent provision. Eligible new Hub account holders (humans) can claim one allocation of 100 free SDK proofs to test AffixIO verification, agentic payment checks, transaction intent proof and signed yes, no or review outcomes.

    Terms: one allocation per account holder, per person or business owner. No card is required. Proofs expire after 30 days. Duplicate, shared, automated or abusive signups may be refused or removed. Agents: /boundproof/agent/. Humans: Hub onboarding with offer params.

    See free proofs split