Caution
Research Dataset — Not for Direct Production Blocking
Auto-aggregated collection of phishing/scam domains from 13+ public threat intelligence sources. Intended for research, pattern analysis, and ML model training.
📌 Purpose
- Aggregate domains from multiple community feeds
- Analyze domain abuse and phishing patterns at scale
- Train AI/ML models and heuristic detection systems
- No operational impact — being listed here doesn't block anything
📂 Files
| File | Description |
|---|---|
blocklist.json |
Merged & normalized domains |
blocklist.txt |
Plain text version |
live_blocklist.json |
DNS-verified (resolvable) |
live_blocklist.txt |
Plain text version |
content_live.json |
HTTP content-verified domains |
content_live.txt |
Plain text version |
dead_blocklist.json |
Non-resolving domains (research) |
state.json |
Per-source hash & counts |
count.json |
Total count badge |
live_count.json |
Live domain count badge |
content_active_count.json |
Content-verified count badge |
🚦 Policy
- No manual removals — fully automated, edits get overwritten
- To remove a domain → report to the original source feed
- For production use → prefer the curated
list.jsoninstead
🛠️ How It's Built
Generated by smart_aggregator.py:
- Fetches from all source feeds
- Normalizes and deduplicates
- Outputs
blocklist.json+ metadata
DNS validation runs separately on the server via dns/active_domains.py:
python smart_aggregator.py
python dns/active_domains.py --force
📡 Sources
| Source | URL |
|---|---|
| MetaMask | eth-phishing-detect |
| ScamSniffer | scam-database |
| SEAL (Security Alliance) | blocklists |
| Polkadot JS | phishing |
| OpenPhish | public_feed |
| Crypto Firewall | crypto-firewall |
| Enkrypt | phishing-detect |
| SPMedia | Crypto-Scam-Threat-Intel |
| Codeesura | Anti-phishing-extension |
| Discord Phishing (Nikolai) | Discord phishing domain feed |
| Discord Phishing (Dogino) | Discord phishing domain feed |
| Phishunt | phishunt.io |
| PhishDestroy | Primary curated list |