Back to GDELT Cloud
Last updated July 22, 2026

GDELTCloudBot

The identity our services present when they fetch web content, and how to control that access. This is the page our fetch user-agent points to.

How to identify us

When GDELT Cloud fetches web content, our requests carry the user-agent string: Mozilla/5.0 (compatible; GDELTCloudBot/1.0; +https://gdeltcloud.com/bot). We identify ourselves honestly and do not disguise our requests as an ordinary browser to bypass access controls.

Fetches originate from our cloud infrastructure. If you need a stable network identifier to allow-list or block, contact us and we will share current details.

What we fetch, and why

We fetch two kinds of content. During ingest, we retrieve news article pages that have been surfaced by our discovery sources, and extract their text so we can code the underlying events into structured records. Separately, our research agent retrieves pages relevant to a specific user question at the time that question is asked.

Our product is derived structured data — coded events, entities, and scores — not a re-hosting of article text. We do not republish full article bodies.

How we try to be a good citizen

We fetch at a modest rate and cache aggressively so we do not re-request the same content unnecessarily. We honor a source's declared content policy: sources that decline automated access are not fetched for their content.

Machine-readable crawl controls are honored on the news-source fetch path — the one that reads publisher feeds and sitemaps. Before each request we read your robots.txt and apply the rules for GDELTCloudBot (falling back to the wildcard group), including Allow/Disallow precedence and wildcard and end-anchored patterns. We wait out any Crawl-delay you set rather than crawling faster, and if we cannot fit the wait into the run we skip the source instead. If robots.txt cannot be served at all we treat that as a no rather than assuming a yes. A text-and-data-mining reservation at /.well-known/tdmrep.json withdraws permission to mine your article text: we may still record that an article exists, but we do not process its body.

This covers the path where we go to publishers directly. Other fetch paths — notably third-party article extraction — are being brought under the same checks. If you have set a control and believe we are not honoring it, tell us and we will prioritize your domain.

Opting out

The fastest way to opt a domain out is to email us; we maintain a per-domain exclusion list and will add you promptly. You can also disallow GDELTCloudBot in your robots.txt.

Opting out removes your content from our fetch paths going forward. It does not by itself remove derived records that were previously coded; ask us if you would like those addressed as well.

Opt out or ask a question

To request that we stop fetching a domain, or to ask about how we access your content, email us and we will action it.