Cohere Web Crawlers
We do not use Cohere bots or user agents for the purpose of crawling or scraping web content to train generative AI foundation models at this time. Should Cohere make use of these tools to crawl or scrape web content to train AI models in the future, Cohere will identify those bots and/or user agents in the table below alongside information about how to disallow their access to your site.
Robots.txt
It is Cohere’s policy to require that crawlers be designed to respect robots.txt. Were we to
use a bot or user agent to crawl or scrape web content to train AI models in the future, web site
operators could follow the following guidance to block the Cohere bot or user agent from accessing
their website:
To block a Cohere bot or user agent from accessing your entire website, add the following to your
robots.txt file in the top-level directory for each sub-domain you wish to exclude.
Example:
Other methods, such as blocking IP addresses, may not work reliably or ensure persistent opt-out, as
they prevent the reading of a site’s robots.txt file.
Updates
If we make updates to this page, we will publish them in the changelog on this site. You can subscribe to an RSS feed to receive updates automatically here.