Key takeaways
- GA4 hides data in two different ways. Thresholding removes whole rows when user counts are low, mainly when Google Signals or demographic data is involved. Sampling approximates data once an Exploration query exceeds 10 million events. Small publishers feel thresholding most.
- GA4 caps granular Exploration data at 14 months on standard properties, with a 2-month default. Aggregated reports last longer, but true event-level history beyond 14 months needs BigQuery export or GA4 360.
- Privacy-first replacement trackers are simpler and cookieless. Plausible starts around $9/mo, Fathom around $15/mo, and Matomo is free self-hosted or around EUR29/mo on cloud. They cost money and hold less behavioural depth than GA4.
- You do not have to remove GA4 to get a cleaner read. A read-only layer over your existing GA4, such as Ramprt, adds no second tracking script, so it will not double-count or dent your ad RPMs, and you keep all your history.
- GA4 undercounts AI traffic. Referrer-less AI visits from in-app browsers and copy-paste links land in Direct, so the true scale of AI discovery stays hidden unless you configure GA4 for it or read it through a purpose-built tool.
If you run an ad-funded site and you are frustrated with GA4, the honest answer is this. You have three real paths: keep GA4 but read it through a cleaner layer, add a privacy-first tracker such as Plausible or Fathom alongside or instead of it, or export raw data to BigQuery. GA4 is free, tightly wired into Google Ads and Search Console, and since May 2026 it natively flags some AI-assistant traffic. But it also thresholds and samples your data, caps granular history at 14 months, and misattributes much of your AI traffic into Direct. Which route is right depends on your traffic size, your budget, and whether your bigger problem is missing data or a messy read of the data you already have. This guide walks through the trade-offs plainly, with no hype.
Why do publishers look for a GA4 alternative in the first place?
Three limitations drive most of the frustration, and it helps to name them precisely because each has a different fix.
Data thresholding. GA4 withholds rows entirely when there are not enough users behind them. There is no fixed public number. Google's own documentation states that data thresholds are "system defined" and cannot be adjusted, and that GA4 withholds rows containing demographic, interest, search-query or Google Signals data when user counts are low. Google deliberately keeps the exact minimums undisclosed. If you run a small or mid-sized publisher site, this hits you far harder than it hits a large brand, and it is the single most common reason publishers say they cannot trust GA4's granular reports.
Sampling. This is a separate mechanism that publishers often confuse with thresholding. Sampling applies to Exploration reports, not standard reports, which are always unsampled. It kicks in when a single Exploration query exceeds 10 million events on a standard property (up to 1 billion on GA4 360), at which point GA4 returns a scaled approximation rather than the full dataset. Thresholding removes rows and gives you nothing in their place; sampling keeps the shape of the data but estimates the numbers. Because thresholding is mostly triggered by Google Signals and demographic data, turning Google Signals off removes most thresholds, at the cost of cross-device and demographic reporting.
Retention. GA4 standard properties retain user- and event-level Exploration data for a maximum of 14 months, with a default of just 2 months. GA4 360 can go up to 50 months. Aggregated standard reports such as Traffic Acquisition and Pages and Screens persist longer, but any granular or custom Exploration analysis is capped at 14 months unless you export the raw data elsewhere. For a publisher trying to compare this year's seasonal performance to two years ago at the page level, that cap bites.
The two data-limiting systems at a glance
| Mechanism | What it does | What triggers it | How to reduce it |
|---|---|---|---|
| Thresholding | Removes whole rows, no estimate given | Demographic, interest, search-query or Google Signals data with low user counts | Disable Google Signals (loses cross-device and demographics) |
| Sampling | Returns a scaled approximation of the data | An Exploration query exceeding 10 million events (standard reports are unsampled) | Narrow the query, use standard reports, or export to BigQuery |
BigQuery export is the honest way past all of this. It bypasses thresholding, sampling and cardinality limits, and gives you raw event-level data with no cap. The catch is that it needs a Google Cloud project, SQL skills and metered query and storage costs. That is why many publishers do not treat raw BigQuery as their "alternative" at all. They treat a simpler dashboard layer as the practical answer, which brings us to the options.
Do I need to remove GA4 to use a GA4 alternative?
No, and for most publishers you should not. There are two very different kinds of "alternative", and conflating them causes needless re-tagging.
The first kind is a replacement tracker: Plausible, Fathom or Matomo. These install their own snippet and collect their own data going forward. You can run one alongside GA4 or switch to it entirely, but it starts counting from the day you add it, so it holds no history of your past traffic.
The second kind is a read-only layer over your existing GA4. A tool like Ramprt reads your current GA4 property through the Analytics Data API, read-only, and presents it more cleanly. It adds no second tracking script, so there is nothing new to tag, no risk of double-counting, and no new data-collection consent question to answer. You keep every year of GA4 history you already have and simply get a better read of it. If your problem is that GA4 is confusing and buries your AI traffic, rather than that GA4 does not collect the data you need, this is usually the lighter fix.
Can a second analytics tracker hurt my ad revenue?
It can, and this matters more for Raptive and Mediavine publishers than for most sites. Every extra script you load competes for the browser's attention and can slow the page, and page speed feeds directly into ad viewability and therefore revenue. More specifically, layering a second full analytics container onto an ad-managed site can create attribution conflicts that muddy your numbers and, in practice, dent RPMs.
This is the strongest practical argument for the read-only route. A layer that reads your existing GA4 via the API adds no client-side script at all, so it cannot slow your pages or interfere with the ad stack. If you do decide to add a privacy-first tracker, the good news is that Plausible, Fathom and Matomo are deliberately lightweight and cookieless, which limits the performance cost compared with a heavy tag-manager setup. But "no new script" will always beat "a light new script" when ad viewability is on the line. If you are weighing this up, our guide to why a second analytics script can hurt ad revenue goes deeper on the mechanics.
Which GA4 alternative is cheapest for a small publisher?
If you want the cheapest paid tracker, Plausible is the usual entry point at around $9 a month for 10,000 pageviews. If you self-host Matomo, the software itself is free forever and you pay only for your own server, which can be the cheapest option of all for the technically confident, though "free" here means you carry the hosting and maintenance. Here is the honest comparison of the main privacy-first alternatives.
| Tool | Starting price | Entry pageview allowance | Best for |
|---|---|---|---|
| Plausible | from about $9/mo | 10k pageviews | Simplest, cookieless, EU-hosted |
| Fathom | from about $15/mo | up to 100k pageviews | Best value once past Plausible's base tier |
| Matomo | free self-hosted, or cloud from about EUR29/mo | varies by cloud plan | Most feature-rich, GDPR capable, on-premise option |
Pricing and allowances are drawn from current 2026 analytics pricing comparisons; check each vendor for the latest tiers. A worked example: a food blog doing 60,000 pageviews a month would sit outside Plausible's base tier but comfortably inside Fathom's up-to-100k plan at around $15, which for that traffic band tends to be the best value of the paid trackers. A hobby site under 10,000 pageviews is cheapest on Plausible's entry plan. And a publisher who simply wants a cleaner read of the GA4 data they already collect can use a free read-only viewer and pay nothing at all, because there is no new tracking to host.
Do not forget the hidden operational cost that favours all three replacement trackers: they are cookieless and generally need no consent banner. Plausible is EU-hosted with no cookies and no personal data, Fathom markets EU isolation routing, and Matomo offers a cookie-free mode. For a publisher with an EU audience, sidestepping the consent-banner overhead that GA4 carries is a real and recurring saving, not just a one-off.
Do GA4 alternatives work with Raptive and Mediavine?
Yes. None of these tools conflicts with an ad network in the sense of breaking it, because analytics and ad management are separate systems. Raptive and Mediavine care about page performance and viewability, not which analytics tool you run. The real question is not "does it work" but "does it slow the page or muddy attribution", which is the ad-revenue point above.
Practically, that means the ranking for an ad-funded publisher is: a read-only GA4 layer is the safest because it adds no script, a single lightweight cookieless tracker is next, and stacking multiple full analytics containers is the thing to avoid. If you already run GA4 for its Google Ads and Search Console integrations, and you probably should keep it for those, then adding a heavy second stack is the worst of both worlds. Keep GA4, and either read it more cleanly or add at most one lightweight tracker.
What can a GA4 reader show that GA4 does not?
The clearest answer for publishers in 2026 is AI traffic. GA4 misattributes a large share of it, and a purpose-built reader can surface what GA4 leaves buried.
The referrer problem. Many AI browsers and apps strip the referer header. Visits from ChatGPT's Atlas browser, in-app browsers and copy-paste links arrive with no referrer, so GA4 files them under Direct or (not set). Perplexity's Comet browser, by contrast, tends to show up cleanly as source perplexity.ai with medium referral. The upshot, explained well in MarTech's breakdown of how GA4 records AI browser traffic, is that default GA4 undercounts AI as a source and inflates Direct.
What Google fixed, and what it did not. Around 14 May 2026, GA4 added a native "AI Assistant" channel to its Default Channel Group, auto-tagging recognised AI referrers with the medium ai-assistant and no configuration needed. Google named ChatGPT, Gemini and Claude as examples at launch but has not published the full list of recognised referrers. This is a genuine improvement, covered in Search Engine Journal's report. But it is purely referrer-based, so any AI visit without a referrer header still falls into Direct. It does not solve the referrer-less problem; it only labels the traffic that already arrives identifiable.
What you still have to build by hand. For engines the native channel does not cover, or if you want tighter control, you have to build a Custom Channel Group in GA4 yourself. Go to Admin, Data display, Channel groups, create a custom group, add a condition group where Source matches a regex such as chatgpt.com|chat.openai.com|perplexity.ai|claude.ai|gemini.google.com|copilot.microsoft.com, then move that channel above Referral and save. It works, but it is manual and brittle, and you must maintain the regex as new engines appear.
A reader built for this does that work for you and goes further than a channel label. It can show AI referrals by engine, which AI crawlers are taking your content, how often AI answers are built on your pages, and what AI-referred traffic earns compared with your site average, without you touching a regex.
Why measuring AI traffic accurately is worth the effort
AI referrals are still a small share of most sites' traffic, often well under a few percent of sessions. But two things make it worth measuring precisely rather than letting GA4 bury it in Direct. First, it can convert unusually well: Similarweb's panel data has ChatGPT referral traffic converting at about 7.1%, second only to paid search at 7.8%. Second, it is growing fast, and you cannot manage a growing channel you cannot see.
There is a caveat worth stating plainly, because honest measurement means honest caveats. Those conversion figures come from third-party panels and vendor studies, not from GA4 itself, and they describe conversion propensity, not traffic volume. AI is still a small share of your sessions. The point of measuring it is to catch a trend early, not to overstate today's numbers.
The bigger picture: what AI is doing to publisher traffic
The reason publishers care about AI measurement at all is the search-side revenue threat behind it, and here the evidence is strong. The Pew Research Center study published on 22 July 2025, based on 68,879 unique Google searches from 900 US adults, found that when an AI summary was present users clicked a traditional result in only 8% of visits, versus 15% without a summary, roughly half. Links inside the AI summary itself were clicked in just 1% of visits, and 26% of AI-summary searches ended the browsing session entirely, against 16% without. That is the click decoupling publishers are trying to see in their own analytics.
The supply side is starker. Cloudflare's crawl-to-refer data quantifies the grievance. As of July 2025, Anthropic's crawlers made roughly 38,000 requests for every referral sent back, a large improvement from about 287,000 to 1 in January 2025, while OpenAI ran about 1,091 to 1, down from 1,217 to 1. Traditional Googlebot sat far lower, roughly 4 to 5 to 1 over the same period. Perplexity, by contrast, moved the wrong way, worsening its ratio by more than 250%. AI crawlers take far more content than they send back in traffic, and that imbalance is exactly what publishers want their analytics to make visible.
So which option should you choose?
Match the fix to the problem. If your issue is that GA4 collects data you cannot get to, because of thresholding, sampling or the 14-month cap, and you have the SQL skills, BigQuery export is the honest raw-data answer. If your issue is that GA4 is complex and you want a simpler, cookieless, consent-light tracker, and you are willing to pay and to start your history afresh, a privacy-first tool such as Plausible, Fathom or Matomo fits. And if your issue is that GA4 already holds your data but reads badly and hides your AI traffic, a read-only layer over your existing GA4 gives you a cleaner view with no re-tagging, no new script, no ad-revenue risk and every year of history intact.
You can see the read-only approach in practice, reading a live demo GA4 with an AI tab that surfaces exactly the traffic GA4 usually hides. Try the free Ramprt demo and read your own GA4 without adding a single tag.
Frequently asked questions
Is GA4 being shut down, so I have to switch?
No. GA4 remains Google's active analytics product and is free. The reasons publishers look elsewhere are its data thresholding, sampling above 10 million events per Exploration query, a 14-month cap on granular Exploration data, and its habit of filing referrer-less AI traffic under Direct. None of that forces a switch, and many publishers keep GA4 and simply read it more cleanly.
Will adding a second analytics tool break my Raptive or Mediavine ads?
It will not break them, because analytics and ad management are separate systems. The risk is indirect: an extra client-side script can slow the page and hurt ad viewability, and a second full analytics container can muddy attribution. A read-only layer that reads your existing GA4 via the API adds no script, so it avoids that risk entirely.
What is the cheapest GA4 alternative for a small site?
Among paid trackers, Plausible starts at about $9 a month for 10,000 pageviews. Self-hosted Matomo is free apart from your own server costs. Fathom is better value once you pass Plausible's base tier, at about $15 a month for up to 100,000 pageviews. A read-only viewer over your existing GA4 can be free because there is no new tracking to host.
Does GA4 now track AI traffic properly on its own?
Only partly. Since around 14 May 2026 GA4 has a native AI Assistant channel that auto-tags recognised AI referrers, with ChatGPT, Gemini and Claude named as examples. But it only works when a referrer header is present, so referrer-less AI visits from in-app browsers and copy-paste links still land in Direct, and engines outside Google's recognised list still need a manual custom channel group.
Do I lose my history if I switch tools?
With a replacement tracker like Plausible, Fathom or Matomo, yes: it starts counting from the day you install it and holds none of your past traffic. With a read-only layer over your existing GA4, no: it reads the data GA4 already holds, so you keep all your history and add nothing new to the site.
Sources
- Google Analytics Help — About data thresholds
- Google Analytics Help — Data retention
- Google Analytics Help — About data sampling
- Search Engine Journal — Google Analytics Adds AI Assistant As Default Channel Group
- MarTech — How GA4 records traffic from Perplexity Comet and ChatGPT Atlas
- Pew Research Center — Google users are less likely to click links when an AI summary appears
- Cloudflare blog — The crawl-to-click gap: AI bots, training, and referrals
- Similarweb — Gen AI Stats 2026
- StackScored — Web Analytics Pricing 2026 (Plausible vs Fathom vs Matomo)