The useful question is not whether AI is fashionable, but whether a defined workflow improves under the organisation’s quality, safety, privacy, cost, and worker-impact constraints.
Tools and models matter, but so do use-case selection, process design, training, governance, worker participation, and measurement. This playbook makes those decisions explicit; it does not promise adoption or productivity gains.
Treat this as a candidate playbook for teams to adapt and evaluate. Organisation size alone does not establish that a rollout pattern will work.
Do not start with licences. Start with a supportable number of bounded workflows, named owners, affected-worker participation, baseline measurements, data rules, and explicit stop conditions.
Evidence anchors: the NIST AI Risk Management Framework treats governance, mapping, measurement, and management as continuing organisational work; the ILO’s 2025 generative AI and jobs update and 2026 review of empirical evidence on jobs, productivity, and work organization emphasize job transformation and implementation context. Neither source validates the example timelines or adoption targets below; those are planning examples to replace with your baseline and staff consultation.
Discover actual practice before designing the rollout
Some staff may already use approved or unapproved AI tools; others may have valid reasons not to. Survey actual use, data handling, accessibility needs, useful practices, failures, and concerns without turning discovery into employee surveillance. Do not infer an adoption gap or competence level from anecdotes.
The playbook matters when the team needs to turn a candidate tool into a documented, governed workflow and determine whether the result is genuinely better.
The first decision: scope
Before any rollout, decide what you’re trying to do. The framings differ:
Productivity uplift. Make existing work faster and better. Each person spends less time on the same outputs. Net: same outputs, fewer hours, or more outputs in the same hours.
Quality uplift. Make existing work better. Net: same hours, higher quality output.
Cost reduction. Reduce headcount, contractor spend, or vendor spend. Net: same outputs, fewer people.
Capability expansion. Do things the team couldn’t do before. Net: new outputs that weren’t possible.
These are different programs with different success criteria. A “productivity uplift” program tracks time saved. A “cost reduction” program tracks headcount or spend changes. A “capability expansion” program tracks new outputs.
If stakeholders want several outcomes, name a primary outcome and record the trade-offs. Otherwise a quality loss can be hidden behind a time-saving claim, or increased workload can be hidden behind higher output volume.
Choose the primary goal from the organisation’s actual problem and stakeholder impacts. “Productivity” is not automatically easiest to measure: time estimates can hide correction work, work intensification, shifted effort, quality loss, or increased monitoring.
Picking use cases
A broad objective such as “use AI in marketing” cannot be evaluated as a workflow. Define the before-state, candidate after-state, affected people, input and output boundaries, and a stop condition.
A useful template:
Use case: [specific task]
Before: [how the team does this today, with concrete time]
After: [how the team will do this with AI, with concrete time]
Owner: [one person]
Decision date: [when we evaluate]
Success criterion: [what would make us declare success]
A hypothetical measurable use case looks like this; its numbers are examples, not expected results:
Use case: Drafting first-pass account research before a sales call.
Before: SDR spends 20-30 minutes per call doing manual research.
After: AI produces a draft in 60 seconds; SDR reviews and adds personal notes in 5 minutes.
Owner: Sales operations lead.
Decision date: 4 weeks from start.
Success criterion: 50% of SDR team uses the workflow weekly; average prep time drops from 25 min to 8 min.
A bad use case looks like:
Use case: Use AI to improve our sales process.
Before: We sell things.
After: We sell things better with AI.
Owner: VP Sales.
Decision date: We'll see.
Success criterion: Increased revenue.
The specific version makes the assumptions and owner reviewable. The vague version does not supply a falsifiable outcome or decision date.
Start with a small number of use cases that the available owners and reviewers can evaluate. The correct number depends on capacity and risk.
The companion template linked from this article gives you the exact fields to use for each candidate use case.
Picking the right first use cases
Some characteristics of a good first use case:
Concentrated time investment. Pick a task multiple people on the team spend significant time on. A workflow saving 30 minutes/week per person across 20 people is 10 hours/week of impact.
Clear input and output. Tasks with well-defined inputs and outputs are easier to automate than ambiguous ones. “Summarise this customer call” is well-defined. “Make our customer experience better” is not.
Low downside risk. Pick tasks where errors are recoverable, not catastrophic. Internal documents over customer-facing emails. Drafts over final outputs. Recommendations over decisions.
Existing measurement. Existing baselines can reduce setup work, but verify that they capture correction effort, quality, worker impact, and shifted work—not only output volume.
Named ownership. Assign someone with time and authority to coordinate the evaluation, collect staff feedback, and stop the workflow when criteria fail. Enthusiasm is useful but does not replace worker participation or accountable ownership.
Don’t pick:
- Use cases where AI is not actually better than current tools.
- Use cases that touch sensitive data without privacy approval in place.
- Use cases with major regulatory implications until you’ve cleared legal.
- “Innovation theatre” use cases nobody actually wants.
Building the workflows
For each use case, the artifact is a workflow — a specific, repeatable process the team uses. Not a vague “use ChatGPT to help.”
A workflow includes:
- The trigger. When does the workflow start?
- The tools. Which AI tool, which model, which integration?
- The prompts. The exact prompts to use. (For consumer tools, the conversation starter. For custom apps, the system prompt.)
- The inputs. What does the human provide?
- The outputs. What does the AI produce?
- The review. Who reviews the AI’s output before it’s used?
- The success metric. How do we know this is working?
Document this in a shared place — wiki, Notion, Confluence. The workflow should be specific enough that a new team member could execute it.
Add one more field: the stop condition. If the output is wrong, if required data is missing, if the task touches sensitive data, or if the confidence threshold is not met, where does the workflow go? Good workflows define the happy path and the refusal/escalation path.
A candidate pattern is for the owner and a representative, consenting pilot group to build and test the workflow before a wider decision. Choose the group and duration from the affected roles, task frequency, risk, and sample required—not from a universal timetable.
The training problem
Generic tool instruction does not establish that staff can perform a specific workflow safely. Training should cover the real task, data boundary, review criteria, refusal path, and incident reporting.
Generic product navigation is insufficient when the required competence is executing and reviewing a particular workflow. Training should cover the task and its failure paths, not assume prior knowledge or adoption.
A workflow-specific training candidate is:
- “Here’s the workflow our sales team uses to research accounts. We’re going to do it together for 3 real accounts.”
- “Here’s the workflow our content team uses to draft blog post outlines. We’re going to do it for 3 real posts.”
- “Here’s the workflow our support team uses to draft replies. We’re going to do it for 3 real tickets.”
Hands-on, on real work, with the specific prompts and tools they’ll use day-to-day.
An illustrative format:
- Day 1: 60-minute workshop. Walk through the workflow on real examples. Each person tries it.
- Week 1: Each person commits to using the workflow on at least 3 real tasks.
- Week 2: Group review. What worked, what didn’t, what do we change?
- Decision gate: Compare the pilot with the baseline, consult affected staff, and decide whether to hold, revise, expand, or stop. Usage alone is not a release criterion.
Treat this as an example plan, not a validated adoption formula. Replace the dates, task counts, and target with a baseline, accessibility needs, staff consultation, and risk-appropriate evidence.
Policy and guardrails
Before scaling, document the rules, ownership, and escalation path. The policy should address privacy, security, intellectual property, employment impacts, accessibility, output review, and applicable sector rules without framing staff as adversaries.
A baseline policy document includes:
Approved tools. Which AI tools is the team allowed to use for work? (Consumer ChatGPT? Only Teams/Enterprise? Specific apps?)
Approved data types. Define permitted data by system, purpose, role, classification, provider contract, and applicable law. “Public” does not automatically mean unrestricted for collection or reuse; personal or confidential data needs an approved lawful and secure path.
Customer-facing rules. Are AI-generated responses to customers OK? Under what review process? Is disclosure required?
Generated content rules. Can AI-generated content be used for X (marketing, sales, internal)? Is human review required?
Review and accountability. Who reviews AI outputs before they’re used in consequential ways? Who is accountable if something goes wrong?
Logging and audit. What gets logged? Who can access logs? For how long are they kept?
Provider data use. Record whether prompts, outputs, files, feedback, and telemetry may be retained or used to improve models under the exact product, plan, region, and settings. Verify the current contract and controls rather than assuming an enterprise default.
Escalation path. What if someone has a question about whether something is OK?
Policy length follows the risk and organisation. Make the operational rules discoverable, specific, accessible to affected staff, linked to the full governing policies, and owned by someone who can answer questions and update them.

Measuring impact
Impact measurement is easy to distort when time saved, correction work, quality, workload, surveillance, and affected stakeholders are measured separately. Define the baseline and counter-metrics before the pilot.
Three levels of measurement:
Level 1: Adoption. Are people using the workflow? (Usage data from the AI tool, self-reported survey, direct observation.) Easy to measure, but doesn’t prove impact.
Level 2: Time/efficiency. How long do specific tasks take, before vs after? (Time tracking, self-reported, sample observation.) Harder but more meaningful.
Level 3: Output quality and quantity. Has the work product changed? More outputs? Higher quality? Better business metrics? Use qualified output review, existing metrics, and appropriate customer or user evidence; volume alone is not quality.
Select measures that can falsify the claimed benefit and expose harm. Adoption alone is not an outcome; include quality and stakeholder-impact measures whenever the workflow can affect them.
Add a fourth check for safety-sensitive workflows: incident and correction rate. Track how often the AI-assisted workflow produced something that required correction, escalation, or rollback. A workflow that saves time but doubles correction work is not mature yet.
A common mistake: declaring victory at Level 1. “80% of the team is using the workflow!” But did anything actually change? Did the team get more work done? Did quality improve? Did customers notice?
Honest measurement sometimes reveals that the workflow didn’t actually save time, or improved one metric while degrading another. That’s important to know. The goal is real impact, not declared impact.
Failure scenarios to design against
Use these as risk scenarios, not prevalence claims:
Failure 1: Tool-first, not workflow-first. Licences are assigned without named workflows, owners, baselines, or decision criteria, leaving impact unmeasurable.
Failure 2: No accountable local owner or worker participation. A central team announces a program without giving affected teams time, authority, or a route to challenge the workflow.
Failure 3: Skipping measurement. “Of course it’s working, look at how excited everyone is.” Excitement is not impact. Measure.
Failure 4: Unusable policy. Rules do not match real work or provide an approved alternative, so exceptions and unresolved needs become invisible. Consult affected staff and provide a usable request/escalation path; permissiveness is not automatically safer.
Failure 5: Missing boundaries. Confidential data is entered into an unapproved service or an unsupported customer-facing claim is sent because data and review rules were never explicit.
Failure 6: Scope exceeds support capacity. Too many roles or workflows launch before the team can answer questions, review incidents, or compare outcomes. Stage expansion according to evidence and support capacity.
Failure 7: Treating it as one-and-done. Workflows that work in May are stale in November because models change, tools change, the team’s needs change. AI adoption is ongoing, not a project.
A staged plan—set dates from evidence
The following phases are a planning scaffold, not a 90-day performance promise. Set duration from task frequency, consultation needs, legal/security review, sample size, and operational readiness.
Scope and selection.
- Decide your primary goal (productivity, quality, cost, capability).
- Identify only as many specific use cases as the review team can support.
- Assign an accountable owner for each and identify affected stakeholders.
- Get baseline measurements where possible.
Workflow design and test preparation.
- For each use case, the owner and a representative, consenting pilot group design the workflow and tests.
- Document it.
- Test on approved representative work for the duration and sample size defined in the evaluation plan.
Policy and review.
- Write or update the operational policy and link it to full security, privacy, employment, procurement, and sector policies.
- Get sign-off from legal/security/leadership.
- Communicate to the whole team.
Training and bounded pilot.
- Run a workshop per workflow.
- Participation and task selection follow the pilot plan, worker consultation, accessibility needs, and applicable employment rules.
- The owner and escalation contacts are available for questions and incidents.
Refinement.
- Group review: what works, what doesn’t, what’s changing.
- Update workflows based on real-world experience.
- Address adoption barriers.
Measurement and decision.
- Pull data on adoption, time saved, output changes.
- Decide for each use case: scale, refine, or cancel.
- Record the next review date and evidence required for any expansion.
At the decision gate, give each use case one of three outcomes:
| Outcome | Meaning | Next action |
|---|---|---|
| Scale | Clear usage, time/quality gain, acceptable risk | Expand to more users or adjacent workflow |
| Refine | Useful but unreliable, unclear metric, or training gap | Fix prompt/process/tooling and re-test |
| Cancel | No meaningful gain or risk too high | Stop the workflow and document why |
Do not declare the workflow operational until its release criteria, staff consultation, training, policy, support, and rollback requirements are met. Later cycles are not guaranteed to be faster.
A note on individual vs team adoption
Individual practice may differ from official tooling. Discover it through a voluntary, non-punitive process that distinguishes approved work use from personal experimentation.
Where staff choose to share approved practices, assess them against the same data, quality, security, accessibility, and worker-impact criteria. Credit contributors and do not convert voluntary experimentation into an undisclosed performance expectation.
A top-down rollout can miss useful practice, accessibility needs, and hidden risks. Evaluate existing workflows with their users rather than automatically codifying or suppressing them.
What the playbook comes down to
AI adoption changes technology, work design, governance, skills, and sometimes roles. Treat those as one socio-technical decision rather than a software-licensing exercise.
The playbook’s review criteria are:
- Pick a clear primary goal.
- Pick specific, concrete use cases (not vague “use AI in X”).
- Build workflows, not just tool access.
- Train on real work, not abstract tool features.
- Set policy that’s specific and workable.
- Measure honestly at multiple levels.
- Iterate based on what they learn.
- Treat it as ongoing, not one-time.
The teams that fail:
- Roll out tools and hope.
- Have vague goals and vaguer measurement.
- Skip the workflow-building step.
- Train abstractly.
- Have either no policy or unusable policy.
- Declare victory at “adoption” without checking impact.
The playbook is a set of testable decisions, not a six-month outcome promise. Keep workflows that show net benefit under the agreed measures, revise those with fixable gaps, and retire those whose risk or total cost outweighs their value.



