We fetch at a modest rate and cache aggressively so we do not re-request the same content unnecessarily. We honor a source's declared content policy: sources that decline automated access are not fetched for their content.
Machine-readable crawl controls are honored on the news-source fetch path — the one that reads publisher feeds and sitemaps. Before each request we read your robots.txt and apply the rules for GDELTCloudBot (falling back to the wildcard group), including Allow/Disallow precedence and wildcard and end-anchored patterns. We wait out any Crawl-delay you set rather than crawling faster, and if we cannot fit the wait into the run we skip the source instead. If robots.txt cannot be served at all we treat that as a no rather than assuming a yes. A text-and-data-mining reservation at /.well-known/tdmrep.json withdraws permission to mine your article text: we may still record that an article exists, but we do not process its body.
This covers the path where we go to publishers directly. Other fetch paths — notably third-party article extraction — are being brought under the same checks. If you have set a control and believe we are not honoring it, tell us and we will prioritize your domain.