Evidence state: Framework. Methodology framework only. No completed HoopAI benchmark, customer result, model comparison, or empirical performance finding is reported. Illustrative examples are not observed results.
Study status. The benchmark below is a design proposal. No completed HoopAI marketing experiment, campaign result, creative ranking, or product-performance finding exists in this paper.
Its purpose is to test evidence-sensitive marketing work without pretending that taste can be reduced to one universal score.
Define the construct before collecting copy
A benchmark measures what its cases and rubric make visible. If the set rewards fluency alone, it cannot answer whether a draft is accurate, usable, or safe to publish.
The construct here is approved-brief transformation. A system receives a bounded evidence packet and produces a specified artifact for a real channel and review state.
Question under study
Can an assisted marketing workflow turn approved material into channel-ready drafts without adding unsupported claims, dropping required context, or bypassing editorial judgment?
The question applies to the workflow, not only the model response. Source presentation, editing controls, approval, and export behavior are part of what is tested.
Build task families instead of a copy pile
Create separate families for synthesis, outlining, first drafts, adaptation, editing, metadata, and claim review. Each family needs its own success definition.
A landing-page introduction and a paid social variant may share facts but differ in space, audience knowledge, disclosure needs, and permissible calls to action.
Do not mix all families into one leaderboard. Results should show where the workflow helps and where the task specification or evidence remains inadequate.
The brief packet
Each case should contain the audience, channel, objective, approved messages, evidence, restricted phrases, required qualifiers, accessibility needs, and final approver.
Mark every objective claim that requires support. Link it to the exact source passage or label it unavailable, so reviewers can distinguish evidence use from plausible invention.
Add an explicit creative range. Some cases may allow broad reframing; others may require close adaptation. Reviewers cannot judge divergence unless that boundary is written.
Sample constraints deliberately
Use a coverage matrix across task family, funnel stage, channel, audience, source quality, and claim risk. The matrix matters more than a convenient total case count.
Include complete briefs, sparse briefs, internal contradictions, obsolete proof points, missing approvals, and requests that cannot be supported by the supplied evidence.
Create paired cases that differ in one condition. For example, one packet contains approved performance evidence while its pair removes that evidence but keeps the requested claim.
Those pairs test whether the workflow changes behavior when substantiation disappears. They are more diagnostic than unrelated prompts with different topics and styles.
Reserve some sectors, channels, or message patterns for a holdout set. Freeze its content before final prompt selection and document any case removed after inspection.
Use a two-layer rubric
The first layer covers non-negotiable evidence and workflow conditions. The second covers editorial quality within outputs that pass the first layer.
- Evidence layer: source fidelity, claim support, qualifier retention, disclosure, and restricted-content compliance.
- Editorial layer: audience fit, hierarchy, clarity, specificity, channel form, accessibility, and amount of revision required.
A polished claim without a reasonable evidence basis fails the evidence layer. High style marks cannot compensate for that failure in a combined average.
The FTC substantiation and endorsement materials can inform questions for US advertising. Responsible counsel must decide which rules apply to an actual campaign.
Calibrating editorial reviewers
Use reviewers with relevant channel experience and a separate evidence reviewer for higher-risk claims. Give them anchor examples and the same approved source packet.
Review independently before discussion. Record both the score and the edit reason, because disagreement about voice has a different meaning from disagreement about factual support.
Test rubric consistency on a calibration subset. If reviewers cannot apply a dimension reliably, revise the definition or report it as qualitative evidence.
Catalogue failures at the point of origin
Distinguish source misuse, brief omission, invented proof, lost qualifier, audience mismatch, channel violation, inaccessible structure, and approval-state error.
Also track workflow failures: a hidden citation, an edit that breaks provenance, a final draft exported before approval, or a restricted term not surfaced to the reviewer.
When a brief is impossible, the expected output may be a targeted question or refusal. Scoring every abstention as incomplete would reward fabrication.
Analysis without a false universal ranking
Report pass rates for evidence gates and distributions for editorial dimensions. Break results down by task family, source condition, channel, and claim risk.
Show reviewer effort as categorized edits and time to approved artifact. Do not equate a fast first draft with a faster publishing process.
Predefine any comparison between configurations. Record the prompt, model, tools, retrieval state, and benchmark version so a later run can explain changed behavior.
A result can support a narrow decision, such as allowing outline assistance with human claim review. It cannot establish brand impact, conversion, or creative superiority.
Known limits
Editorial preference remains partly subjective even with anchors. A static benchmark can age quickly as channels, audiences, product facts, and advertising rules change.
Source quality can dominate model behavior. English-language cases do not establish performance across languages, cultures, markets, or accessibility contexts.
Offline judgment does not measure campaign lift, incrementality, brand preference, or commercial return. Those outcomes require different designs and live governance.
What a future results report must reveal
A later empirical paper should publish the task taxonomy, sampling logic, benchmark version, exclusions, rubric, reviewer credentials, agreement method, and uncertainty.
It should disclose sponsorship, system access, prompt assistance, manual edits, protocol deviations, and all critical evidence failures, including those later corrected.
This proposed benchmark reports no completed HoopAI run. Its authoritative sources provide design context only and do not endorse the method or any vendor.
Methodology
Create task families for research synthesis, outlining, first-draft generation, adaptation, editing, and metadata. Give every case a reader, channel, approved source packet, required message, restricted terms, claim boundary, call to action, and review state. Include briefs that are complete, ambiguous, contradictory, and intentionally missing evidence. Freeze a representative holdout set and record every model, prompt, context, and tool change.
Score source fidelity and claim safety before style. Then assess audience fit, structural clarity, brand constraints, accessibility, channel requirements, and amount of editing required. Use independent reviewers for high-risk claim dimensions and record why they disagree. Treat a polished draft with an invented capability as a failure even if its tone and structure are strong.
Limitations
- Editorial preference is partly subjective and needs calibrated examples.
- A benchmark can underrepresent new channels, languages, audiences, and policy changes.
- Source quality may influence results more than model choice.
- This framework does not measure campaign lift, conversion, brand preference, or commercial return.
Sources
- FTC Policy Statement Regarding Advertising Substantiation: Primary US advertising-policy source for evidence behind objective claims. It is legal background, not a legal opinion or approval of this protocol.
- FTC Guides Concerning Endorsements and Testimonials in Advertising: Primary US guidance for endorsement evidence and disclosure questions. Applicability must be assessed for the actual campaign and jurisdiction.
- NIST AI 600-1, Generative AI Profile: Public generative-AI risk reference. It provides no evidence of a HoopAI implementation or satisfied control.
Notes
Methodology framework only. No completed HoopAI benchmark, customer result, model comparison, or empirical performance finding is reported. Illustrative examples are not observed results.







