DisputeDB — UDRP Legal Research
Citation Integrity
Each search result includes a pre-formatted citation field. Use it verbatim when citing cases, overview sections, or policy paragraphs. Do not cite UDRP case numbers from memory — LLM training data contains incorrect and fabricated case numbers (e.g. “D2005-0960”) that either do not exist or involve different parties. If no relevant result is found, say so rather than inventing a citation.
Overview
datapulse_disputedb_search provides semantic vector search over a curated UDRP (Uniform Domain-Name Dispute-Resolution Policy) knowledge base. It enables legal research across panel decisions, authoritative guidance, and the foundational policy documents — all through a single query interface.
Database Composition
| Source Type | Description | Chunks | Items |
|---|---|---|---|
decision | WIPO and Forum (NAF) panel decisions (1999–2026) | ~927,700 | ~116,900 cases |
overview_section | WIPO Overview 3.0 — panel consensus views | 101 | 64 sections (1.1–1.15, 2.1–2.15, 3.1–3.12, 4.1–4.22) |
guide | WIPO Guide to the UDRP — Q&A explainers | 52 | 52 items (A01–I01) |
policy | ICANN UDRP Policy (9 sections + sub-paragraphs) | 35 | 35 paragraphs |
policy | ICANN UDRP Rules (21 rules + sub-paragraphs) | 107 | 106 paragraphs (one is split across two chunks) |
Curated chunks: 295 (measured 2026-09-05, after the Overview re-ingest).
Total: ~928,000 chunks, all embedded with BGE-M3 (1024-dimensional vectors) and searchable (measured 2026-09-05). Newly ingested or re-parsed decisions join the index at the next embed run.
Decisions break down as ~73,200 WIPO and ~43,800 Forum (measured 2026-09-05). Roughly 98,900 have full decision text and are therefore searchable; the remaining ~18,000 are metadata-only records for cases that were terminated, withdrawn, or never published in full.
Refresh cadence: Decisions are updated via periodic discovery/ingestion runs. Policy, rules, overview, and guide are static reference documents re-ingested when upstream changes occur.
Chunking Strategy
All source documents are split into overlapping chunks of ~2,000 characters with 200-character overlap. Split points prefer paragraph breaks, then sentence boundaries, and are always rune-aligned (safe for multi-byte UTF-8).
chunk_index=0is the first chunk of a source document- Higher
chunk_indexvalues are later sections of longer documents - Most guide and policy items fit in a single chunk (
chunk_index=0) - Decisions average ~9 chunks each
Search Parameters
| parameter | range | notes |
|---|---|---|
query | 1–1000 characters | required; whitespace-only is refused |
limit | 1–20 | omit or pass 0 for the default of 10; above 20 is capped to 20 |
min_score | 0–1 | default 0.55; 0 disables the filter; outside 0–1 is refused |
source_type | decision, overview_section, guide, policy | omit to search everything |
include_all_cited_cases | boolean | return every cited_cases entry instead of the first 25 |
limit: 0 is the default, not zero rows. The API reads it the same way, so the
two never disagree.
Response Format
datapulse_disputedb_search reports returned — how many results came back.
It has no total, because a similarity search has none: every chunk matches to
some degree and the limit IS the universe. datapulse_disputedb_lookup reports
count as the true total with returned beside it, because a lookup does have
one. Two names for two different things.
The response also echoes min_score (the threshold actually applied) and, when
a source_type filter was passed, source_type. An empty result set carries a
note explaining it in terms of the threshold, so “nothing scored above 0.97”
is not mistaken for “the corpus has nothing on this subject”.
{
"query": "bad faith registration",
"returned": 10,
"min_score": 0.55,
"results": [
{
"id": "uuid",
"source_type": "decision",
"source_id": "uuid",
"chunk_index": 3,
"content": "The Panel finds that the Respondent registered the domain...",
"metadata": {
"case_number": "D2024-0001",
"case_uid": "wipo:D2024-0001",
"source": "wipo",
"outcome": "transfer",
"domain_names": ["example-brand.com"]
},
"score": 0.7142,
"case_number": "D2024-0001",
"case_uid": "wipo:D2024-0001",
"domain_names": ["example-brand.com"],
"outcome": "transfer"
}
]
}
Prefer the top-level decision fields
case_number, case_uid, domain_names and outcome appear both at the top
level of a decision result and inside metadata. Use the top-level fields.
Both are read from the decision record when you query: since 2026-08-26 the
metadata block is derived from the chunk’s parent at query time rather than
stored with the chunk, so the two no longer diverge (the stored copies had
drifted on a third of decisions). The top-level fields remain the documented
contract and are the ones to read.
Metadata Schemas by Source Type
decision
| Field | Type | Example | Description |
|---|---|---|---|
case_number | string | D2024-0001 / 1733616 | Provider’s own case number. WIPO uses D<year>-<seq>; Forum uses a bare integer that carries no provider marker. |
case_uid | string | wipo:D2024-0001 | Provider-qualified identifier, unambiguous across providers. Split on the first colon — WIPO and ADNDRC case numbers contain hyphens, so a hyphen cannot be the separator. |
source | string | wipo / forum | Provider |
outcome | string | transfer | transfer, denied, cancelled, terminated, withdrawn, split decision; case active marks a case still pending |
domain_names | string[] | ["example.com"] | Disputed domain(s) only — see below |
domain_names contains disputed domains, not the complainant’s
This distinction matters and is easy to get backwards. domain_names lists the
domains the case was brought against. It does not list the complainant’s own
domains.
So a brand owner’s primary domain will not appear here, because nobody ever
disputed it. Searching domain_names for microsoft.com returns nothing; that
is correct, not missing data. UDRP decisions routinely cite a complainant’s own
domain as evidence of trademark rights, and such mentions are deliberately
excluded.
This tool cannot answer “which disputes did X file”
That question is keyed on complainant, a decision field that is not carried in
chunk metadata. datapulse_disputedb_search searches decision TEXT and returns
chunks, so it has no way to filter by who brought a case.
The two fields describe opposite relationships and are easy to confuse:
| question | field | tool |
|---|---|---|
| was this domain attacked? | domain_names | either |
| does this owner enforce? | complainant | datapulse_disputedb_lookup |
Do not infer enforcement history from domain_names. A brand’s own domain is
absent from it precisely because nobody disputed it.
Use datapulse_disputedb_lookup with complainant: for that question. It reads
complainant records only — not respondents, not disputed domains — so the answer
is not contaminated by cases where the name appears on the other side.
(Complainant coverage is all but 2 of ~73,200 WIPO decisions, and ~91% of Forum, where it is recovered from the provider’s case caption rather than supplied as a field.)
Some decisions carry fused domain entries
A residue of rows carry several domains concatenated into a single string — about 15 domain rows on 2026-09-05, down from ~3,600 after the provider listings were re-read and the fallback path repaired:
"domain_names": ["nexusautopartsusa.comnexusautopartsus.com"]
"domain_names": ["dragons-softwares.comdragons-stores.comdragons-supports.com"]
The parser dropped the <br> separating them. Both paths are fixed and their
rows repaired; what remains are values the conservative detection rule cannot
separate with confidence. For those, a domain lookup on one of the
constituent names misses that case.
A lookup marks these where it can: a domain object carries malformed: true
when the value is several domains run together.
{"seq": 0, "domain": "nexusautopartsusa.comnexusautopartsus.com",
"source": "delimited", "malformed": true}
Two things about that flag. source is PROVENANCE, not a quality claim — a
fused value genuinely came from the provider’s listing, so source: "provider"
says nothing about whether the value is intact, which is why malformed says it
separately. And detection is deliberately conservative: it fires on a repeated
TLD or three or more markers, because counting markers alone would flag
foo.commerce.net. An absent flag is not proof a value is clean —
dolphins.comjets.com is a syntactically valid domain name and nothing
distinguishes it with certainty.
Where the flag is absent and an entry still shows a TLD in a non-final position, cite the case number and omit the domains rather than quoting a fabricated one.
overview_section
| Field | Type | Example | Description |
|---|---|---|---|
section_number | string | 2.1 | Overview section (1.1–4.22; 64 in all) |
title | string | "2.1 Identical..." | Section heading |
source | string | wipo_overview | Always wipo_overview |
element | string | first | UDRP element: first, second, third, procedural |
cited_cases | string[] | ["D2000-0003", ...] | Case numbers the section cites; the first 25 by default |
cited_cases_total | int | 1030 | Only on a cut list: how many the section really cites |
cited_cases_truncated | bool | true | Only on a cut list |
cited_cases is a hand-curated citation graph, not an extraction. It answers
“what are the leading cases on X” deterministically, without semantic search:
retrieve the Overview section for the point of law, then look up the cases it
cites. Section 1.7 carries 28 of them.
The list is cut to its first 25 entries by default, and a cut list says so
with cited_cases_total and cited_cases_truncated: true. Pass
include_all_cited_cases: true to receive every entry. Since the 2026-09-05
re-ingest the median section cites 21 cases and the largest, 3.8, cites 113
(1,521 citations across the 61 sections that carry any), so the cap bites on
20 sections and never by much. Before it, a parsing bug attributed
neighbouring sections’ lists to one section (1,107 on 1.11) and one result ran
to 20 KB of case numbers; the cap holds whatever upstream sends, and the
marker fields mean a cut is never mistaken for the whole.
guide
| Field | Type | Example | Description |
|---|---|---|---|
question_id | string | A04 | Question identifier (letter + number) |
question | string | "What is bad faith?" | Full question text |
section | string | A | Section letter (A–I) |
section_heading | string | "Scope of the UDRP" | Section title |
source | string | wipo_guide | Always wipo_guide |
policy
| Field | Type | Example | Description |
|---|---|---|---|
document | string | udrp_policy | udrp_policy or udrp_rules |
paragraph | string | 4(a)(i) | Paragraph number with sub-sections |
title | string | "a. Applicable Disputes" | Paragraph title (if any) |
source | string | icann | Always icann |
Score Interpretation
Scores are BGE-M3 cosine similarity (0–1) between the query embedding and chunk embedding. BGE-M3 produces lower absolute scores than some other models: on this corpus the ceiling is roughly 0.75–0.80, and only paraphrase-style queries reach ~0.81, so treat 0.75 as a guide rather than a bound.
| Score Range | Meaning | Action |
|---|---|---|
| 0.72+ | Strong semantic match | Directly relevant — cite with confidence |
| 0.65–0.72 | Relevant | Likely useful context — read and evaluate |
| 0.55–0.65 | Weak match | May contain tangentially relevant info |
| Below 0.55 | Poor match | Unlikely to be useful (filtered by default) |
Guide Sections (A–I)
| Section | Topic | Items |
|---|---|---|
| A | Scope of the UDRP | 9 |
| B | Overview of the UDRP | 5 |
| C | Preparing and Filing a Complaint | 9 |
| D | Preparing and Filing a Response | 13 |
| E | Role of the Panel | 4 |
| F | Panel Decision | 8 |
| G | Role of the Registrar | 1 |
| H | Role of WIPO | 2 |
| I | Resource Materials | 1 |
Key Policy Paragraphs
| Paragraph | Topic |
|---|---|
| 4(a) | Three elements complainant must prove |
| 4(a)(i) | Identical or confusingly similar |
| 4(a)(ii) | No rights or legitimate interests |
| 4(a)(iii) | Bad faith registration and use |
| 4(b) | Non-exhaustive bad faith factors |
| 4(c) | Demonstrating legitimate interests |
| 4(k) | Availability of court proceedings |
Workflow Patterns
Standalone Legal Research
disputedb_search("What constitutes bad faith in UDRP?", limit=15)
→ Mix of decisions, overview sections, guide items, and policy paragraphs
→ Cite: [D2024-0001], [Overview 3.0, Section 3.1], [UDRP Policy, Para. 4(b)]
Domain Dispute Analysis (combine with DNS tools)
1. searchlabels("brandname") → find lookalike labels + TLD spread
2. live_rdap("brand-name-fake.com") → registration details
3. disputedb_search("confusing similarity with trademark for <domain>")
4. disputedb_search("respondent bad faith parking page with ads")
→ Build case analysis with precedents + registration evidence
Policy Lookup
Pass source_type: ["policy","overview_section","guide"]. The 295 curated
chunks compete with ~928,000 decision chunks in one index, so a question phrased
differently from the curated heading is usually answered by decisions QUOTING
the rule rather than by the rule itself. Filtering makes this deterministic.
disputedb_search("UDRP rules response filing deadline", limit=5)
→ Returns UDRP Rules paragraphs on response timing
Three-Element Analysis
1. disputedb_search("identical or confusingly similar <trademark> <domain>")
2. disputedb_search("rights or legitimate interests <respondent's argument>")
3. disputedb_search("bad faith registration and use <specific circumstances>")
→ Structured analysis matching panel decision patterns
The two disputedb tools
| tool | answers |
|---|---|
datapulse_disputedb_search | reasoning — bad faith, legitimate interests, panel consensus, policy. Semantic search over decision text. |
datapulse_disputedb_lookup | facts — a specific case, a specific domain, or what a company has filed. Structured lookup. |
Reach for lookup when you already know the thing you are asking about.
Semantic search cannot reliably answer “has anyone disputed example.com”,
because it depends on that string appearing in some chunk’s text.
datapulse_disputedb_lookup
Exactly one of case_number, domain, complainant, respondent, or panelist.
case_number: "D2000-0003" one decision; retired identifiers still resolve
domain: "example.com" disputes brought AGAINST that domain
complainant: "Microsoft" what that company has filed
respondent: "Fundacion …" what was brought against them
panelist: "Neil Anthony Brown" cases they DECIDED
Each of the three party modes reads its own records. A name appearing on the other side of a dispute does not contaminate the answer.
domain and complainant are opposite relationships. A brand’s own domain
is absent from disputed-domain lists precisely because nobody disputed it, so
enforcement history is only answerable through complainant.
Truncation is explicit. List results are capped (default 20, max 100;
limit: 0 is the default, not one row; a row is ~1.3 KB) and the
response carries the true count alongside returned, plus truncated and a
note when rows exist beyond the page. The busiest complainant in the corpus
has 687 decisions and Microsoft 444, so a capped answer is common — say it is
partial rather than presenting it as a full history.
Paging. complainant, respondent and panelist lookups take offset
(default 0, max 100,000): rows are ordered decision_date DESC then
case_number DESC, so pages are stable, and the note names the next offset —
offset: 0, limit: 20 is rows 1–20, offset: 20 is rows 21–40. truncated
means rows exist beyond THIS page (count > offset + returned), and the
response echoes offset. A page past the end says so rather than “no
matches”. case_number and domain lookups ignore offset with a note: one
returns a single decision and no domain has more than a handful of disputes.
domain is an EXACT match; complainant, respondent and panelist match
by words. Every whitespace-separated word of the query must appear in the
recorded name, in any order, each as a case-insensitive substring, so
panelist: "Debrett Lyons" finds “Debrett G. Lyons” and “Mr. Debrett Gordon
Lyons” alike. A partial word still matches (Lyon finds Lyons); there is no
stemming. Matching is literal — no fuzzy or phonetic matching, so Newman
never finds Neuman; on zero results retry with the surname alone or a shorter
distinctive word before concluding the name is absent. This asymmetry is
deliberate and is not visible from the parameter names. A bare
label is rejected — domain: "equifax" is refused as a single label, because a
label is not a domain name — while complainant: "apple" correctly matches
Apple Inc., Snapple and Applebee’s. Case and trailing root dots are normalized
away before the lookup, and internationalized names are folded to punycode, so
Tesla.com and tesla.com. are the same lookup and instagrạm.com and
xn--instagrm-tx0d.com reach the same row. A leading www. is sent as part of
the name — the response’s query reads domain:www.tesla.com — and the
database matches it away itself, so www.tesla.com finds the same decisions as
tesla.com. When normalizing changed more than letter case, the response
carries submitted_domain with what you sent. See
datapulse_help(topic="normalization").
Party lookups need at least 3 characters, counted over the whole query rather
than per word (G. Lyons is accepted; G. alone is not). complainant: "HP" is refused
with an error naming the floor rather than answered with an arbitrary sample,
because a two-character substring matches most of the corpus. Use a longer
distinctive form such as HP Inc.
A case number that does not exist returns count: 0 with a note, not an
error. Treat it as “no such case”, not as a transient failure to retry.
Sort order is decision_date DESC, then case_number DESC. A capped list
is therefore the most recent matches, which makes “X’s recent filings”
answerable — but still say the list is partial.
Parties are lists, and each carries its own provenance. A dispute names as many complainants and respondents as it names:
"respondents": [
{"seq": 0, "name": "Carolina Rodrigues", "source": "delimited"},
{"seq": 1, "name": "Fundacion Comercio Electronico", "source": "delimited"}
]
seq is the provider’s listed order — the first complainant is the lead filer.
source says how that specific value was obtained:
source | meaning |
|---|---|
provider | the provider stated it — authoritative |
structural | separated from the provider’s own listing markup |
delimited | separated from a provider string by a known delimiter |
llm | a model separated a value whose delimiter was lost — a reconstruction |
kind is reserved for classifying a party as entity, privacy_service,
registrar or placeholder, but no row carries it yet (measured
2026-09-02): the column exists and nothing populates it. An empty kind means
unclassified, not “not a real party”, and filtering on it today drops
everything. Respondent lists routinely include privacy services and registrars
alongside the actual registrant; they are kept rather than discarded, so read
the names rather than relying on kind.
An empty complainants list means “could not be determined”, NOT “filed
nothing”. A parties_note is attached where this matters.
panel — who decided the case
panel lists the panelists, one row each with seq. They are the
adjudicators, not parties, which is why they have their own lookup mode and
why the general decision search does not match them.
It is present on about 83% of decisions (96,826 of 116,942 on 2026-09-05, up
from 73% three days earlier after signature-block parsing landed upstream).
Metadata-only cases — terminated, withdrawn — never had a panel, and roughly
2,100 decided cases carry a panel line the separator could not read. An absent
panel therefore means “not recorded”, not “no panel”.
A UDRP panel is one member or three (93,575 and 2,548 decisions respectively). Anything else is a parsing artifact rather than a real panel — the corpus still holds 647 two-member “panels” that are one name split in the wrong place, and 56 with more than three.
Counting by panelist name undercounts badly. There is no entity resolution: one panelist is written many ways, and Neil Anthony Brown alone has over 200 spellings in the corpus (“The Honourable Neil Anthony Brown QC”, “Hon. Neil Brown”, “the Hon Neil Brown Q.C.”). A word search finds them; grouping or totalling by exact name does not. Report what the search returns and say it is a lower bound.
Decision text is omitted by default — the median decision is ~14,600 characters.
include_text: true returns it for a single case_number lookup, truncated to
8,000 characters. The response then always says what happened:
text_available: false with a note for the ~19,000 metadata-only decisions,
text_truncated: true with text_chars when it was cut. On a domain,
complainant, respondent or panelist lookup the flag is ignored and the note says so.
Errors and Retries
A request DisputeDB never answered is reported with the endpoint, the network
error, the attempt count and the time spent, e.g. DisputeDB unreachable: POST /search: dial tcp 127.0.0.1:8081: connect: connection refused (2 attempts, 251ms). A connection-level failure — nothing listening, or a pooled connection
the server had already closed — is retried once automatically before that error
is returned, so the caller need not retry it again.
A request the server answered is never retried. A 4xx names the mistake to
correct (an unknown source_type, a party query under 3 characters); a 5xx
such as embedding failed is an upstream outage that will not clear within a
call.
DisputeDB request timed out: POST /search gave no response in 30s means the
server accepted the connection and did not answer. One retry by the caller is
reasonable; repeated retries are not.
Related Topics
datapulse_help(topic="searchlabels")— Find lookalike/typosquatting domainsdatapulse_help(topic="typo_squatting")— Typosquatting detection workflowdatapulse_help(topic="rdap")— Domain registration detailsdatapulse_help(topic="reconnaissance")— Full domain investigation
Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026. Your AI assistant reads this page by calling datapulse_help(topic="disputedb").