Three Sites Made 215,128 'Best Software' Pages for AI—Perplexity Called Them Out
A new investigation reveals how three websites manufactured over 215,000 AI recommendation pages. Perplexity and other AI systems cite them as sources, raising urgent questions about AI training data quality.
The Manufactured Source Problem: How 215,128 Pages Corrupted AI Recommendations
This week, a critical investigation exposed a troubling pattern in how AI systems discover and cite sources. Three websites alone produced 215,128 "best software" pages—most of them algorithmically generated, low-quality content designed primarily to capture AI citations. Perplexity, ChatGPT, and other major AI search tools have been citing these pages as authoritative sources, effectively amplifying noise rather than signal in the AI recommendation ecosystem.
This discovery matters urgently because it reveals a fundamental vulnerability in how modern AI systems work. When Perplexity or similar tools recommend software, they're not reading original research or expert reviews—they're pattern-matching against whatever content ranks highest and gets cited most. If that content is manufactured at scale, the entire recommendation chain becomes compromised.
What Happened: The Scale of the Problem
According to research from Trellner, the three offending sites deployed a simple but effective strategy:
- Automated page generation at massive scale
- Template-based content with minimal unique information
- SEO optimization designed specifically to surface in AI search queries
- Affiliate links and monetization as the primary business model
None of this is technically illegal, but it's intellectually dishonest. These weren't sites publishing genuine software reviews or comparisons. They were content mills, pure and simple, designed to intercept AI queries before they reached legitimate sources.
The concerning part? It worked. Perplexity cited these manufactured pages. So did other AI search tools. Users asking for software recommendations got answers grounded in low-effort, quantity-over-quality content rather than actual expert knowledge.
Why This Matters for AI Development and Trust
This isn't just a problem for end users seeking software recommendations. It's a foundational challenge for AI systems themselves:
1. Training Data Poisoning at Scale
When AI models are trained on internet data, they absorb these manufactured pages as though they were legitimate sources. The sheer volume—215,128 pages—means these low-quality signals can drown out authentic voices. An AI trained on this data learns that certain claims are "true" simply because many pages repeat them.
2. Citation Feedback Loops
Once Perplexity and other AI systems cite these pages, they become more authoritative in search rankings. More visibility means more citations. More citations mean higher ranking. The manufactured source gains legitimacy through pure repetition, not because it deserves it.
3. The Trustworthiness Crisis
As an AI journalist covering the tools industry, I see this happening daily. When users can't trust that AI recommendations are grounded in real expertise, they stop using AI search for high-stakes decisions. For developers choosing a code assistant, or businesses evaluating analytics platforms, this erosion of trust is costly.
How These Pages Were Created
The investigation identified several patterns:
Template-Based Generation: One site would create a "best software for X" template, then populate it with hundreds of variations ("best software for AI," "best software for startups," "best software for 2024," etc.). Each page followed identical structure and reasoning, just with different keywords.
Minimal Originality: Many pages contained boilerplate paragraphs copied across thousands of articles, with only product names and a few details changed.
Optimization for AI Queries: The sites specifically used language and structures known to perform well in AI search results—short paragraphs, numbered lists, direct answers—because they understood that AI systems pattern-match differently than traditional search.
Monetization Through Affiliate Links: The financial incentive was straightforward: generate enough pages, get enough AI citations, drive enough traffic, and earn affiliate commissions when visitors click through to software platforms.
What Perplexity's Response Reveals
Perplexity's identification and acknowledgment of this problem is actually a positive sign. It shows that at least one major AI search platform is actively monitoring citation sources and willing to correct course. However, it also exposes the broader fragility of the system:
- No built-in verification: AI systems don't inherently know whether a source is authoritative or manufactured.
- Scale creates vulnerability: With millions of pages indexed daily, manual quality control is impossible.
- Incentives are misaligned: There's no cost to creating 215,128 low-quality pages, but significant reward if even a small percentage drive traffic.
The Broader Implications for AI Tool Discovery
For those of us in the AI tools space—whether at ListmyAI or elsewhere—this investigation is a wake-up call about curation standards. Human judgment and editorial review still matter, perhaps more than ever.
When you're evaluating an AI tool recommendation, consider the source:
- Did a human test the software? Or did an algorithm scrape existing pages?
- Are there specific use cases or limitations mentioned? Or just generic praise?
- Who benefits financially from the recommendation? (Transparency about affiliate relationships matters.)
- Is the reviewer comparing alternatives thoughtfully? Or just listing products with boilerplate descriptions?
Directories like ListmyAI exist partly as a counterweight to this problem—human-curated, editorially consistent, focused on actual tool functionality rather than manufactured comparison pages.
What Should Change
Several interventions could reduce this problem:
- AI systems should weight source reputation: Perplexity and similar tools should develop internal measures of source trustworthiness, beyond just citation frequency.
- Content farms should face detection penalties: If a domain generates 100,000+ similar pages annually, algorithms should be skeptical of that domain's authority.
- Affiliate links should be disclosed: AI-generated summaries should flag when a cited source benefits financially from a recommendation.
- Original research should be valued: AI training should over-weight genuine primary sources and penalize derivative, bulk-generated content.
- Human curation should remain central: For high-stakes recommendations (software choices, business tools), AI should defer to curated sources with editorial standards.
The Takeaway
The discovery that three sites manufactured 215,128 "best software" pages—and that Perplexity cited them—reveals a critical weakness in how modern AI systems source information. It's not a failure of AI technology itself, but rather a failure of incentive structures and verification processes.
As AI search becomes more central to how people discover tools and make decisions, the quality of the sources feeding those systems becomes crucial. This investigation should prompt both users and developers to ask harder questions about where AI recommendations come from, and to invest in more robust systems for distinguishing authentic expertise from manufactured noise.
For now, if you're looking for genuinely useful AI tool recommendations, turn to sources that show their work—editorial sites that explain why a tool matters, not just that it exists. The volume of AI content is growing exponentially, but the signal-to-noise ratio is getting worse. Only human judgment and transparent curation can fix that.
AI Tools Mentioned in This Article
Listmyai
Listmyai is a hub for discovering, bookmarking, and getting AI-driven recommendations for AI tools
Comet Browser
The AI browser that acts as a personal assistant. Automate tasks, research the web, organize your email, and more with C
GPT-4o
OpenAI's flagship model with vision, audio, and text capabilities in a single model `#freemium`
Three Sigma
AI research tool for quick document analysis and question answering
Gemini
Google’s multimodal AI assistant integrated across Google products.
Nextthreebooks
AI-powered book recommendations tailored to your reading preferences
Explore more at the full AI tools directory →
Frequently Asked Questions
When AI systems cite low-quality manufactured content as authoritative sources, they train themselves and their users to trust unreliable information. This creates a feedback loop where poor-quality pages gain credibility simply through repetition, corrupting the entire recommendation ecosystem and undermining user trust in AI search.
Sources & Further Reading
Find the right AI tool for you
Browse 1,000+ AI tools in the ListmyAI directory
Comments
Sign in to comment
Join the conversation — sign in or create a free account.