AnyLearn
All lessons
Businessintermediate

Measuring Whether It Worked, and Protecting the Brand

Marketing's usual metrics cannot answer whether AI helped, because output volume rose and attribution is already hard. This lesson covers what to measure instead, the brand and legal risks that concentrate in this function, disclosure norms with clients and audiences, and the honest reckoning on which claimed gains survive scrutiny.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 8

Why the usual metrics cannot answer this

Ask a marketing team whether AI helped and the answer arrives quickly: we produce four times the content. That is a real change and it is not an answer.

The problem is that output is an input, and marketing's job is outcomes. Four times the content with the same distribution and the same attention produces four times the content, not four times the result. In a channel with fixed attention it can produce less, since undifferentiated volume performs worse per item.

And the outcome metrics that would settle it are exactly the ones already contaminated. Conversions, pipeline and revenue move for reasons unrelated to how the copy was produced, and marketing's attribution problem predates AI entirely.

So three failure modes recur when teams report on this.

Measuring activity and calling it impact. Pieces published, variants generated, hours saved. All real, none of them evidence that anything improved.

Attributing a good quarter to the tool. The quarter had other causes, and the tool is the most recent change, which is not the same as the cause.

And not measuring at all, on the reasoning that everyone is doing it. That is a decision to spend without knowing, which is precisely the position the function exists to help others avoid.

What can actually be measured is narrower and more useful, and it is the next step.

Full lesson text

All 8 steps on one page, for reading, reference, and search.

Show

1. Why the usual metrics cannot answer this

Ask a marketing team whether AI helped and the answer arrives quickly: we produce four times the content. That is a real change and it is not an answer.

The problem is that output is an input, and marketing's job is outcomes. Four times the content with the same distribution and the same attention produces four times the content, not four times the result. In a channel with fixed attention it can produce less, since undifferentiated volume performs worse per item.

And the outcome metrics that would settle it are exactly the ones already contaminated. Conversions, pipeline and revenue move for reasons unrelated to how the copy was produced, and marketing's attribution problem predates AI entirely.

So three failure modes recur when teams report on this.

Measuring activity and calling it impact. Pieces published, variants generated, hours saved. All real, none of them evidence that anything improved.

Attributing a good quarter to the tool. The quarter had other causes, and the tool is the most recent change, which is not the same as the cause.

And not measuring at all, on the reasoning that everyone is doing it. That is a decision to spend without knowing, which is precisely the position the function exists to help others avoid.

What can actually be measured is narrower and more useful, and it is the next step.

2. What can actually be measured

Three tiers, in descending order of how much they establish.

Controlled comparison is the only tier that establishes causation. Where the work is testable, run AI-produced and human-produced variants against each other in a properly powered test. This directly answers whether the output performs, and creative testing is the one part of marketing where this is routinely available.

Process metrics establish that the work changed. Time from brief to publication, cost per piece, and the number of variants a test could afford. These are honest and they measure efficiency rather than effect.

Quality proxies establish whether the output holds a standard. Rework rate, the proportion of drafts requiring substantial editing, and claim errors caught in review. A rising rework rate means the time saved in generation is being spent in editing.

And the counterweights, which belong beside any reported gain: rework, claim errors reaching publication, audience response, and whether differentiation is eroding.

The honest report says the process got faster by this much, the output performed the same or better in a test, and these were the costs. Most reports state the first and omit the rest.

flowchart TD
A["Did AI help?"] --> B["Controlled comparison: AI vs human variants, powered test"]
A --> C["Process metrics: time to publish, cost per piece, variants affordable"]
A --> D["Quality proxies: rework rate, claim errors caught"]
B --> E["Establishes causation, where testable"]
C --> F["Establishes the work changed, not that outcomes improved"]
D --> G["Establishes the standard held"]
E --> H["Report all three plus counterweights"]
F --> H
G --> H

3. The counterweights

Every efficiency claim in this function should be reported alongside four numbers, and their absence is usually not deliberate.

Rework rate. What proportion of generated drafts require substantial editing rather than light touch-up. If generation saved two hours and editing added ninety minutes, the gain is half what was claimed, and that arithmetic is rarely done.

Claim errors reaching publication. The count of factual assertions that were wrong and got through. This is the metric that matters most and the one nobody tracks, because catching it requires someone to go back and check published work rather than reviewing at the point of publication.

Per-item performance. Not total results, which rise with volume, but performance per piece. If output tripled and per-piece engagement halved, total engagement rose by half while the work tripled, which is a worse position disguised as growth.

And differentiation. Harder to measure and worth attempting: is your messaging becoming more similar to competitors'? A simple version is periodically comparing your copy against theirs and asking whether a reader could tell them apart.

The reason these get omitted is structural rather than dishonest. The gain is visible immediately and lands with the team reporting it; the costs are delayed, diffuse, and land elsewhere. Which is exactly why the reporting rule has to be procedural: counterweights in the same document, always.

4. Brand risk concentrates here

Marketing publishes, which means its errors are public in a way most functions' are not, and the failure modes are specific.

The fabricated claim reaching a customer. Covered in the first lesson and it remains the largest exposure, because it is a legal and a brand problem simultaneously.

The tonal misfire. A model has no awareness of context beyond its prompt, so scheduled or automated content published during an unfolding event can be badly wrong in a way that is obvious to any human. This is not new, and generation at volume increases the number of opportunities.

The borrowed voice. Output that closely resembles a competitor's distinctive phrasing, or a recognisable creative property, because the model learned from a corpus containing it. Legally this is territory worth taking seriously, and reputationally it undermines the differentiation argument entirely.

Generated imagery containing artefacts or implausibilities that a casual reviewer misses and an audience does not. Public attention to this is high, and the reaction is disproportionate to the actual harm.

And the disclosure failure, where an audience discovers content was generated after the fact rather than being told. The discovery is worse than the fact, which is the pattern the disclosure norms in the managing-teams cursus describe.

The operational response is unglamorous: a kill switch for scheduled content, a human check before anything publishes during a sensitive period, and a record of what was generated so a question can be answered.

5. Intellectual property, briefly and honestly

The intellectual property position around generated content is genuinely unsettled, and the honest treatment is to say what is known and what is not rather than to assert either comfort or alarm.

What is reasonably clear. Copyright protection generally requires human authorship in most major jurisdictions, so purely machine-generated output may not attract protection. For marketing this matters where you need to prevent others copying an asset, since an unprotectable asset is one a competitor may reuse.

What is unsettled. Whether training on copyrighted material without licence is permitted, which is being litigated in several jurisdictions with outcomes pending and inconsistent so far. And the extent of liability for output that resembles training material.

What organisations actually do in the meantime. Human involvement in creation, which improves the protection argument and is good practice regardless. Provider indemnities, which several major vendors now offer for enterprise customers and which are worth reading carefully rather than assuming. And avoiding generated content for assets intended to be distinctive brand property, such as a logo or a signature campaign concept, where protectability is the point.

The practical instruction: for high-value distinctive assets, involve humans substantially and keep a record of that involvement. For volume content where protectability does not matter, the question is less consequential.

And check your vendor's indemnity terms, because they vary and the difference is material.

6. Disclosure, to audiences and to clients

Two disclosure questions, with different answers.

To audiences. The Article 50 duties from 2 August 2026 require disclosure of deepfake content and of AI interaction, and machine-readable marking of synthetic outputs. Beyond the legal minimum, the principle from the managing-teams cursus applies: disclose where it changes how someone should treat the content. Nobody needs a note that a model helped draft a product description. Synthetic imagery of a person, a generated testimonial, or an AI presenter is a different matter, and audiences have reacted badly to discovering these late.

To clients, for agencies, this is the sharper question and it is frequently unresolved. A client paying for creative work has an interest in how it was produced, and the range of current practice runs from full disclosure through silence. Three positions worth deciding between deliberately: disclose the use, disclose on request, or contractually address it at the outset.

The consideration that decides it: if a client would be unhappy to learn later, the arrangement is unstable regardless of whether disclosure was required. And pricing follows from this, since an agency billing on time while generating output has a conversation coming that is better had early.

The general position across both: the discovery is worse than the fact, and every disclosure decision should be made against that asymmetry rather than against the minimum obligation.

7. The skills question

A marketing function adopting AI changes what its people need to be good at, and the shift is worth naming because it affects hiring and development.

What becomes less scarce. Producing competent copy at volume. Adapting content across formats. First-pass translation. These were skills people were hired for and their scarcity has fallen.

What becomes more valuable. Judgement about what to say, which is the brief, and which the first lesson identifies as the stage that determines quality. Editing critically, particularly the ability to detect a plausible false claim, which is a different skill from editing for style. Understanding the customer well enough to know when generated output is subtly wrong about them. And measurement literacy, since the function is now awash in output that needs evaluating.

The development implication that mirrors the managing-teams cursus. The tasks that built copywriting judgement, drafting many pieces and being corrected, are the tasks now automated. So a junior marketer's development path needs deliberate replacement, which in practice means reviewing generated work critically against a standard, taught explicitly rather than absorbed.

And the uncomfortable observation for the function: a team whose distinctive value was producing polished copy has had that value commoditised, and the response is to move up the stack toward strategy, positioning and customer understanding rather than to produce more polished copy faster.

8. The honest reckoning

Pulling the cursus together into claims that survive scrutiny.

The efficiency gains are real and specific. Variant generation, channel adaptation, localisation first passes, metadata at scale, and research synthesis. These are measurable through process metrics and, in the testing case, through outcomes.

The quality gains are unproven and mostly unmeasured. Where teams claim better content, the evidence is usually more content, and per-item performance frequently fell.

The strategic risk is genuine. Cheap generation for everyone means convergent messaging, and a function whose differentiation erodes has lost something the efficiency does not replace.

The legal exposure concentrates in two places: unsubstantiated claims, and targeting in regulated categories including the recruitment advertising most marketing teams do not realise is high-risk.

And the highest-value use is the least discussed. Research synthesis improves what you decide to say, which matters more than how fast you say it.

The closing instruction is the one this function is best placed to accept, because it is what marketing tells everyone else: measure the outcome rather than the activity, report the counterweights beside the gain, and be suspicious of a claim that arrived without a test. A marketing team that applied its own standards of evidence to its own AI adoption would be ahead of most.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. Why is 'we produce four times the content' not an answer to whether AI helped?
    • Content volume cannot be measured reliably
    • Output is an input, and with the same distribution and attention it produces more content rather than more results
    • The comparison period was too short
    • Content volume is confidential
  2. Which tier of measurement actually establishes causation?
    • Process metrics such as time from brief to publication
    • Quality proxies such as rework rate
    • A controlled comparison of AI-produced against human-produced variants in a powered test
    • Year-on-year revenue comparison
  3. Output tripled and per-piece engagement halved. What actually happened?
    • Total engagement rose by half while the work tripled, which is a worse position disguised as growth
    • Total engagement tripled
    • The measurement is invalid
    • Per-piece engagement is not a meaningful metric
  4. What is the reasonably clear part of the intellectual property position?
    • Training on copyrighted material without licence is settled law
    • Generated content is always protectable if edited
    • Provider indemnities cover all output-related claims
    • Copyright generally requires human authorship, so purely machine-generated output may not attract protection
  5. What is the asymmetry that should govern every disclosure decision?
    • The discovery is worse than the fact
    • Legal minimums exceed audience expectations
    • Clients care less than audiences
    • Disclosure reduces engagement measurably

Related lessons

AI
intermediate

What a Percentage Does and Does Not License

A model went from 27 percent to around 57 percent, so it is more than halfway to AGI and the rest arrives shortly. That inference is wrong in at least four ways, and working through why is more useful than the score itself. This lesson covers the linearity assumption, construct validity, contamination, and what the framework is good for once you stop reading it as a progress bar.

8 steps·~12 min
Law & Compliance
advanced

Proof: Disclosure, Presumptions, and the Complexity Rule

Strict liability is worthless if the claimant cannot prove a defect they never saw. Articles 9 and 10 answer that with a disclosure order, three presumptions of defectiveness, a presumption of causation, and a rule turning complexity into the claimant's ally. This lesson works through the cascade, the three-year and ten-year clocks, and what a defendant should be able to produce.

10 steps·~15 min
Law & Compliance
advanced

Who Pays, and For What Damage

The Directive builds a chain of liable operators so an injured person in the EU always has someone to sue. This lesson covers the manufacturer and component manufacturer, the importer and fulfilment service provider route, the distributor's one-month rule, online platforms, how a modification makes you a manufacturer, the heads of damage including data loss, and the exemptions.

10 steps·~15 min
Law & Compliance
advanced

Defectiveness: The Safety a Person Is Entitled to Expect

A product is defective when it lacks the safety a person is entitled to expect. Article 7 turns that into circumstances a court weighs, several written for software: the ability to learn after release, interconnection, cybersecurity requirements, and recalls. This lesson works through the list, the rule that a later improvement is not an admission, and why compliance is not a defence.

10 steps·~15 min