Nexum
We unlock human expertise for the AI economy
Brands we’ve helped implement AI:

Built to make AI work for you.
We're an agency that skips the complexity, focusing only on automation that delivers real, measurable results.
Hours Saved Monthly
Reduction In Manual Work
Time To Measurable ROI
From confusion to clarity, fully handled.
Getting real results without the guesswork.

01
Discovery & Audit
Week 1
We analyze your workflows and pinpoint where automation delivers the most measurable value.

02
Design & Launch
Weeks 2-4
We build tailored automation systems and roll them out carefully without interrupting daily operations.

03
Train & Optimize
Ongoing
We train your team and continuously refine the system to keep results improving.
Real Businesses, Real AI Results
Benchmark 01 · Procurement
Frontier models do the paperwork. They miss the judgment.
Our first benchmark dataset puts AI agents inside a simulated company and hands them real procurement work: renewals, vendors in trouble, pilots turning into production deals. Every task, trap and grading rule comes from a working procurement practitioner.
Each task runs in a live environment with an inbox, purchase requests, contracts, an org chart, approval policies, and colleagues and vendors who answer back. The agent has to decide who to talk to, in what order, and what not to do. We grade the whole path, not just the final answer: did it involve the right people, follow the approval chain, keep internal information away from the vendor, and skip work that didn’t need doing.
Head-to-head score
Same tasks, same grading. 1.00 is how a practitioner would handle them.
Claude Opus 5
0.66
Gemini 3.1 Pro
0.23
Attempts with a firing-level mistake
e.g. approving spend over the limit without the CFO, or cancelling a critical platform with no replacement lined up
Claude Opus 5
20%
Gemini 3.1 Pro
60%
Where Claude Opus 5 goes wrong
Share of attempts where the mistake was possible
Went to the CFO before Finance had reviewed the spend
100%
Re-ran a credit check on a vendor already under contract
100%
Called a 12-month paid commitment a “pilot”
100%
Escalated to the C-suite when nobody asked
53%
0%
attempts either model got fully right
87%
Opus 5 attempts that missed a step the practitioner calls essential
3x
how much more often Gemini made a firing-level mistake
A generic AI grader rated some of these attempts as good as or better than the expert solution, and it praised the extra credit checks and executive updates. The practitioner marked those same moves as mistakes. That gap is what Fieldwork measures.
Both models ran as agents with access only to the environment’s tools. The head-to-head is one attempt per task per model. The mistake chart is Claude Opus 5 across several variants of each task, so answers can’t be memorised. Criteria are weighted by how much the practitioner says each one matters. More models coming soon.
This is what makes us actually different.

With Nexum
Fast Turnaround, No Delays
One Point Of Contact
Custom-Built For Your Business
Regular, Honest Updates
No Hidden Costs
Without Nexum
Long Waits For Fixes
Multiple Vendors, No Accountability
Templates That Miss The Mark
Updates Come Rarely
Costs That Creep Up
VS











