We look at the impact data-scraping robots from AI firms are having on the online encyclopedia used by hundreds of millions of people. Also in this edition of Tech Life: if you work in the fashion industry, if you are a fashion model, are you worried about AI ? A lot are, and we find out why. And how do we prevent children from seeing online adult content ? Many parts of the world are requiring sites to verify the ages of their users. Now the biggest adult site argues that we need a better system. <...
Transcript
Automatically generated from the audio. May contain errors.
This BBC podcast is supported by ads outside the UK. Or listen to the global story on BBC.com or wherever you get your podcasts. Hi and welcome to Tech Life on the BBC World Service, the programme about technology and how it's changing all our lives.
I'm Chris Valance. This week we look at the impact data-slurping robots from AI firms are having on the online encyclopedia hundreds of millions of us visit. I am, of course, talking about Wikipedia.
If you work in the fashion industry, if you are a fashion model, are you worried about AI? A lot are, and we'll be finding out why. And how do we keep children from seeing adult content?
Many parts of the world are requiring sites to check the ages of their users. Now the world's biggest adult site argues we need a better system. Now when you have a question, where do you turn to for answers? Wikipedia, the free online encyclopedia, gets billions of visits every year.
It's available in a wide range of global languages and amazingly it's all produced by unpaid volunteer editors. Lots of us turn to it for information and so too do the AI companies who send out programs called bots to scrape information from the site, perhaps to help their systems to answer
our questions. Recently, Wikipedia discovered a large amount of traffic. It had thought where human readers was, in fact, AI company bots apparently coming from Brazil.
And when they dug into the numbers, it also appeared its human readership had significantly declined. Perhaps because more of us are turning to artificial intelligence to answer our questions. Selina Deckelman is the Chief Product and Technology Officer for the Wikimedia
Foundation. I asked her what was going on. Well, what we've observed is over the last few months, since about May, we've seen a decline in human traffic.
And we think it's about 8%, which is significant for us. And we believe it's primarily driven by automation. So bots that are seeking pages from our website and then using those to do things
like train artificial intelligence or feed search engines. And is it that those answers, the answers that have created by generative AI systems are sort of getting in the way of people going to Wikipedia. So they're sort of, you know, they're taking your content, if you like, these sort of programs are scraping your
content. And that content is being used by generative AI to create articles. And that means that people don't need to go back to Wikipedia. Well, because we can't observe what every person does on their individual devices. We don't know exactly, but what we have seen out in the world from reports from other publishers,
industry peers, we know that other publishers like newspapers and websites, they're seeing declines in traffic as well. So we had anticipated that we would start to see a decline to Wikipedia. The other thing that we know is over the last several years,
people have started using other kinds of media to get their news. So for example, short form media on a variety of sources.
So like TikTok, YouTube, these types of things. People are using that short form video as well. So while we can say that it's mainly, we might think that it's just AI.
I don't think it's just that. We suspect that it's like a variety of things happening all at the same time. You didn't notice this at first.
At first you thought your traffic hadn't fallen, but then you uncovered it if you like, and that was all to do with some disguised Brazilian bots. You'd better tell me about that.
Yeah. So we have these systems that monitor how much traffic we're getting, and they tell us things like where we think
the traffic's coming from. And what we noticed in May was that there was a big spike in unusual but human seeming traffic. So we investigated that and as we dug deeper,
we realized that this traffic was actually automated. When we say bots, that's what we mean is that there's some program that someone wrote and it's going in and like scraping data.
So as we dug into that, what we realized is that our systems for detecting bots, they needed to change a little bit. So that's what we did.
We changed our system for detecting automated traffic and then we looked back at all of the data that we had since May. And that is when we saw and detected this overall decline.
And it's a, the way that we look at it as we compare year to year. So compared to the traffic that we were seeing in 2024, we're seeing about an 8% decline.
And we should be clear, these programs are basically sort of scraping the contents of your site. In the case of AI firms, that would be in order to get information to train their AI system or for their AI system to produce answers with.
Yeah, that's right. You know, your listeners may know that there are a lot of companies right now that have gotten funding to go and try to create the coolest new AI tool, know that consumers will want to use. So there's a little bit of, I would say, a gold rush in this regard, right?
People are trying to create these new tools. They're trying to get data as quickly as they can. And in that work, they're not necessarily thinking about what effect that might have on the sites that they're scraping.
So something that we've been doing for Wikipedia and all the Wikimedia projects is where we created a system for companies to actually work with us, call it enterprise. And what we're trying to do is redirect some of that automated traffic over there so that we protect the services that humans use.
People don't think about these things because the internet is just there. you know, there's a lot of computing that goes into keeping your site, keeping Wikipedia one of the most popular sites on the internet going and serving all the all the demands that are put upon it by human beings, but increasingly by these robots. And that costs you money.
In fact, the type of requests they're making you wrote said that it costs you more money than other types of reading of your content, if you like. That's right. And I don't want to get
This is the opening of the episode. Open the player for the full interactive transcript with clickable words, translation and flashcards.
Open full transcriptAudio belongs to its publisher and is played from their feed. Rights holders can request removal — copyright & takedown policy
Episodes · Tech Life

The future of flying?

Understanding AI Agents

Smell the future

Too young to scroll?

Viva technology!

Microsoft's big quantum bet

Teaching in the AI world





