DNS Neighborhood (datapulse_dns_neighborhood)
Look up DNS infrastructure cluster membership for a domain or cluster ID.
Overview
The datapulse_dns_neighborhood tool reveals which HDBSCAN cluster a domain belongs to and provides cluster metadata plus the exemplar domain (closest to the centroid). All 409M domains are assigned to exactly one cluster: 351 normal clusters or 25 anomaly clusters.
Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
domain | string | No | — | Domain name to look up (e.g., “google.com”). The spelling is canonicalized (case, whitespace, trailing dots, IDN to punycode) and no labels are removed; a single label (localhost) is rejected. When canonicalizing changed more than letter case, the output carries a hint line submitted_domain: <your input> (sent as <name>). See datapulse_help(topic="normalization") |
cluster_id | integer | No | — | Direct cluster ID lookup |
is_anomaly | boolean | No | false | Set true when looking up an anomaly cluster by ID |
At least one of domain or cluster_id must be provided. If both are given, domain takes precedence.
Modes
1. Domain Lookup
Look up which cluster a domain belongs to:
{
"domain": "google.com"
}
Response:
[{
"domain": "google.com",
"cluster_id": 338,
"is_anomaly": false,
"anomaly_score": 0.815,
"nearest_normal_dist": 0.815,
"cluster_name": "awsdns.com*awsdns.org*googledomains.com",
"cluster_notes": null,
"member_count": 135874,
"core_sample_count": 1000,
"radius": 1.069,
"mean_density": 1.342,
"mean_anomaly_score": null,
"max_anomaly_score": null,
"exemplar_domain": "android-anti-virus.com"
}]
2. Normal Cluster Lookup
Look up a normal cluster by ID:
{
"cluster_id": 338
}
Response:
[{
"cluster_id": 338,
"is_anomaly": false,
"cluster_name": "awsdns.com*awsdns.org*googledomains.com",
"cluster_notes": null,
"member_count": 135874,
"core_sample_count": 1000,
"radius": 1.069,
"mean_density": 1.342,
"mean_anomaly_score": null,
"max_anomaly_score": null,
"exemplar_domain": "android-anti-virus.com"
}]
3. Anomaly Cluster Lookup
Look up an anomaly cluster by ID (set is_anomaly to true):
{
"cluster_id": 6,
"is_anomaly": true
}
Response:
[{
"cluster_id": 6,
"is_anomaly": true,
"cluster_name": null,
"cluster_notes": null,
"member_count": 3,
"core_sample_count": null,
"radius": null,
"mean_density": null,
"mean_anomaly_score": 0.687,
"max_anomaly_score": 0.696,
"exemplar_domain": "gptif.xyz"
}]
Response Fields
Domain Lookup Fields
| Field | Description | Normal | Anomaly |
|---|---|---|---|
domain | The queried domain name | Yes | Yes |
cluster_id | Cluster ID (normal or anomaly) | Yes | Yes |
is_anomaly | Whether the domain is in an anomaly cluster | Yes | Yes |
anomaly_score | Distance from cluster center (lower = closer to centroid) | Yes | Yes |
nearest_normal_dist | Distance to nearest normal cluster (equals anomaly_score for normal cluster members, since their nearest normal cluster is their own) | Yes | Yes |
Cluster Metadata Fields
| Field | Description | Normal | Anomaly |
|---|---|---|---|
cluster_name | Label derived from dominant nameserver patterns across the cluster’s membership (does not describe any individual domain’s DNS) | Yes | null |
cluster_notes | Optional annotation | Yes | null |
member_count | Number of domains in the cluster | Yes | Yes |
core_sample_count | Core samples identified by HDBSCAN | Yes | null |
radius | Cluster radius in embedding space | Yes | null |
mean_density | Average point density | Yes | null |
mean_anomaly_score | Average anomaly score | null | Yes |
max_anomaly_score | Maximum anomaly score | null | Yes |
Exemplar Field
| Field | Description |
|---|---|
exemplar_domain | Domain closest to cluster centroid (lowest anomaly_score) — the most representative member |
Concepts
What is an Exemplar?
The exemplar domain is the cluster member with the lowest anomaly_score — meaning it is closest to the cluster centroid in embedding space. It represents the “most typical” domain for that cluster’s infrastructure pattern. Use the exemplar with dptechsim to explore a cluster’s neighborhood.
Normal vs Anomaly Clusters
- Normal clusters (351): Dense groupings of domains with similar DNS infrastructure. These have descriptive names derived from the dominant nameserver patterns across the cluster’s membership, plus metrics like radius, density, and core sample count.
- Anomaly clusters (25): Domains with unusual infrastructure that don’t fit neatly into normal clusters. These have anomaly-specific metrics (mean/max anomaly scores) but no descriptive names.
Every domain is in exactly one cluster — there are no unassigned “noise” points.
Understanding cluster_name
The cluster_name (e.g., awsdns.com*awsdns.org*googledomains.com) is a label derived from the most common nameserver patterns across all members of that cluster. It does not mean every domain in the cluster uses those nameservers. A domain may end up in a cluster due to similarity across all dimensions of its embedding (nameservers, ASNs, MX records, registrar, registration date, TLD) even if its individual nameservers differ from the label.
Understanding nearest_normal_dist
This field is populated for all domains. For normal cluster members, it equals anomaly_score because the nearest normal cluster is their own. For anomaly cluster members, it represents the distance to the closest normal cluster — useful for understanding how far an anomaly domain is from “normal” infrastructure patterns.
Workflows
“Who’s in my neighborhood?”
neighborhood(domain="suspicious.com")
→ cluster_id: 42, exemplar: "typical-domain.com"
→ dptechsim(domain="typical-domain.com", similarity=0.90)
→ See all domains in/near that cluster
Cluster exploration
neighborhood(cluster_id=42)
→ member_count: 12500, cluster_name: "cloudflare*..."
→ dptechsim(domain=exemplar_domain, similarity=0.92)
→ Browse cluster members
Anomaly investigation
neighborhood(domain="odd-domain.com")
→ is_anomaly: true, cluster_id: 6
→ neighborhood(cluster_id=6, is_anomaly=true)
→ See anomaly cluster metadata
Combined with registrar lookup
neighborhood(domain="example.com")
→ cluster_id: 338
→ dptechsim(domain=exemplar_domain)
→ registrar(iana_id=registrar_iana_id) for each match
Relationship to Other Tools
| Tool | Use With neighborhood For |
|---|---|
dptechsim | Explore cluster members via exemplar domain |
searchlabels | Find domains by name, then check their cluster |
registrar | Resolve registrar IDs from techsim results |
dphistory | Track infrastructure changes for cluster members |
Availability
This tool requires the dbserver daemon to be running. If unavailable:
- Tool will not appear in
tools/listresponse - Health check will report
unhealthystatus - Calls fail with the socket connection error, verbatim
Check availability via datapulse_health().
Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026. Your AI assistant reads this page by calling datapulse_help(topic="neighborhood").