At Magellan AI, we have done a deep read and analysis of the AMP Accords. Overall, we believe the AMP Accords are a step in the right direction, and we support the approach.
That said, we think there are details within the proposed documentation that need a wider group to implement, especially at scale. We look forward to participating in larger discussions on the proposals within IAB forums, but in transparency, are also putting our recommendations here for public consumption and any additional comment.
Our largest points of feedback are for control groups, as we think the sole option as laid out disproportionately affects smaller publishers, agencies, ad servers and advertisers, while benefiting companies at scale. Podcasting should continue to be a low entry point medium across the board, so we think more optionality aids in that effort.
As a bigger picture point outside of standardization, podcasting has grown because so many companies, large and small, came together to build it. We want companies of every size and technical capability to be heard.
We’ve broken out our thoughts, support and revisions on each section of the AMP Accords in line with the document.
Definition of a Podcast
Magellan AI fully supports this definition. It is specific enough to distinguish podcasting from adjacent formats and flexible enough to accommodate where the medium is heading. We encourage the industry to keep that flexibility in view as podcasting continues to evolve. We should keep innovating and expanding our offerings while maintaining truth to the underlying shorthand that "if it works with your eyes closed, it's a podcast."
Exposure Metrics
In general, we support these new metrics for both content exposure and advertising exposure. We also support the provision of additional identifiers for consented users.
However, we do think the technical implementation needs to be further refined with further input from podcast players and hosting platforms.
- For non-HLS players:
- This seems to require major updates to how media is delivered and the functionality of the non-HLS players with no input from them.
- This is requiring each player to pass data through an analytics URL within metadata, but it's unclear if the intent is to have one singular analytics URL from the hosting provider or ad server (which seems more reasonable) or potentially multiple for each analytics partner. As we see today in podcasting, one show often has a multitude of prefixes or ad tags appended to each ad. Each added step increases the chances that one URL isn't called or that the functionality is too complicated for the player to work efficiently.
- For hosting platforms:
- Are hosting platforms computing these metrics from the players themselves, and then sharing those with analytics partners? From the documentation itself, it seems as if hosting platforms and analytics partners will all be computing independently.
- If hosting platforms and analytics partners are all computing this independently, even with the same data set, this would seem to lead to the same "lack of uniformity" issue the document raises in its case against players self-computing these metrics. We see this today in various IAB-certified counts, between hosting platforms and analytics partners today.
- Additional concerns:
- Regarding consent frameworks, who will be the responsible party for data collection requests?
- From a legal perspective, are players operating as Data Controllers?
- What does a hosting platform / publisher operate as?
- Who are DPAs to be executed between? Players <> hosting platforms, players <> analytics partners?
- What is the incentive for players to take on this legal risk / effort?
- Suggestions:
- While we agree the suggested method would work, we would suggest a Server-to-Server driven approach as an additional alternative. Being able to chunk delivery and consumption stats over an agreed upon time-frame may be an easier path for certain players to calculate the metrics as laid out, without overwhelming podcast players that are not enterprise-sized businesses.
- The potential to deliver one file with a common schema to multiple parties (hosting platform, analytics partners, etc...) may be an easier development lift, especially in the shorter term as an interim alternative.
Randomized User-Level Holdout Groups
To start, we think that Randomized User-Level Holdout groups could be a useful exercise, but we do not think they should be codified as the only type of incrementality test, and think that other options, which we will expand upon, are able to scale better with less exposure to contamination or impression waste.
- The comparisons brought in this section look specifically at Meta, Google and Amazon — single-source companies of ad delivery and measurement. They control the end-level user connection, ad delivery and the tools that measure those ads. Podcasting works differently. Podcasts are played across many platforms, delivered by various ad servers, and consumed differently even by the same user.
- Within a single ad server, using platform-native user or subscriber IDs, or the user identifiers, would be a workable solution. However, when considering an advertiser buy across multiple platforms and shows, as well as listener behavior, there is no assurance that a user held out on platform A would not be served the ad on platform B. Given the propensity for similar audiences to exist on similar programming this increases the chance for contamination unless platforms are sharing user-data, which would be against the privacy regulations as cited in other sections of this document.
- Baked-in ads break a RUHG in a way better identifiers cannot fix. The read is fixed in the file. Anyone who plays the episode hears it, so the ad server has no way to withhold it from a holdout listener. On a buy that runs both formats, a listener held out of the dynamic ad can still hear the same advertiser baked into another show in the flight. The two groups no longer differ only on ad exposure, which is the condition the document sets for a valid holdout.
- The gap lands where performance is strongest. Our Q1 2026 data put host-read response at 2.45%, ahead of produced at 1.51% and programmatic at 1.93%, and host-read inventory is frequently delivered baked-in. Treating baked-in inventory as an exception case gives podcasting's best-performing ads its weakest measurement.
- The document proposes building holdout groups from listeners of other episodes of the same show, which could cause multiple issues:
- Over a limited period, if an advertiser is running multiple ads in a program, few listeners in that show's audience will have avoided the ad entirely. There may not be enough of them to reach statistical significance.
- If an advertiser is only running one ad and comparing to multiple episodes without an ad, exposure may be too small to produce a material difference.
Here are additional statistically sound alternatives we believe are more easily implemented and scalable:
Synthetic holdout groups
- With the addition of user-level data as laid out in Part Two, analytics providers will be able to construct more defined psychographic and demographic profiles on users that were exposed to ads, for both dynamic and baked-in advertising. Through consented data, analytics providers can create synthetic control groups that mirror the delivered ad audience group and compare performance between the two groups.
- Benefits of this method include:
- Easier set-up for ad servers and advertisers: there is not a net new technological lift.
- Always-on capabilities: testing is not confined to a short window chosen for statistical significance or cost.
- No contamination: with user level identifiers, control groups can be constructed across shows and platforms.
- No wasted inventory: publishers can maximize their show level inventory without losing 20% of it to PSA or ghost bidding.
Same-show consumption-based control groups
- Much of the AMP documentation aims to improve accuracy in consumption metrics. We should be using the newly provided consumption metrics on a show-level to gauge who has and has not heard an advertiser's ad, and act accordingly.
- A current pain point, as laid out, is that many users download without ever consuming. This leads to ad impressions or baked-in downloads being counted that were never consumed. In the case of RUHGs:
- For dynamic ads, the ad server decides whether to deliver the advertiser's ad or a ghost ad at the time of the initial download or stream request. For general podcast delivery (inclusive of HLS and RSS), when someone hits download or play, the entire episode with all ads is constructed at that time.
- A listener served the episode may never listen to the ad if it is a download or if the listener never makes it to that portion of the file. Under the new consumption methodology, publishers may not be paid for that impression. The listener still counts as exposed in the holdout assignment, but never heard the ad.
- In the reverse, a listener that receives the ghost ad / PSA may listen to it, and be in the RUHG. For all intents and purposes, these two listeners could be swapped for greater accuracy, but an ad server cannot know how long a listener plans to listen or if they plan to listen at all at the time of an ad being served.
- Proposal: instead of a RUHG, serve every user the ad and let analytics partners build the holdout groups from lack of consumption, rather than a ghost-ad or PSA. Benefits of this method include:
- Use the new AMP standard to its fullest: instead of adding technical levers that create new failure points, use new consumption stats for greater accuracy.
- No wasted inventory: as admitted in the document itself, sellable overall impressions will likely go down as there is a focus on ads with confirmed listenership. Instead of giving away an additional 20% of impressions (and likely more given consumption or lack thereof), this ensures that publishers can adopt the greater overall standards without losing more inventory.
- Accounts for cross-platform listening: as with synthetic control groups, identifying listeners that may have downloaded or been served an ad but not consumed on more than one platform would not discount them from a control group. Alternatively, if they listened on one podcast but not another, it would ensure that they are part of the exposed group and not contaminating the control group.
We support where the AMP Accords are pointing. Our concern is the single mandated path for control groups. Incrementality testing should be available to a company of any size, and there is more than one sound way to build a holdout.
The next stage of this work happens in the IAB forums, and we will be there with the recommendations above. We would encourage anyone with technical or operational concerns to put them on the record as well. The more companies stress-test this, the more durable it gets.