CrawlSign
Pre-release / design partnersCrawler safety for AI ingestion.
CrawlSign runs inside your crawler and checks every page before it enters a dataset, an index, or an agent's context. It returns one decision per page (continue, log, quarantine, or stop) with a score, stable reason codes, and the evidence behind them.
- What it checks
-
- Refusal signals: robots.txt, X-Robots-Tag, robots meta, noai and noimageai
- Known poison sources and tarpit endpoints
- Hidden-link traps that lead crawlers into generated mazes
- How it runs
- CLI, Python SDK, and a local HTTP API, with adapters for Scrapy, LangChain, and LlamaIndex.
- Where it runs
- On your infrastructure. Page content never leaves your systems.
{
"url": "https://example.com/noai",
"action": "stop",
"score": 115,
"reasons": [
"ai_refusal_signal_detected",
"meta_robots_disallow"
]
}