Investigation Methodology
By Seme Research Team · Updated May 22, 2026
Seme's identity investigation methodology is based on Open Source Intelligence (OSINT) principles, combining AI-driven automated search with structured analysis. This methodology ensures every investigation undergoes systematic multi-source verification, producing auditable and reproducible results. Unlike traditional background checks that rely on single databases, Seme searches across 15-20 independent sources in parallel, confirms facts through cross-validation, and quantifies the credibility of each finding through a five-tier evidence classification system (E1-E5).
Investigation Workflow
Five-Step Investigation Process
Research Planning
AI analyzes input information and designs a multi-angle investigation strategy covering professional, social, academic, and public records dimensions. The system automatically generates 5-8 investigation directions, each with specific search queries and expected data sources.
Multi-Round Search
Parallel execution of SearXNG meta-search engine and OpenAI Web Search API queries, collecting information from 15-20 independent sources. Each round extracts structured facts including names, positions, education, social profiles, and affiliated organizations.
Gap Analysis
AI evaluates current investigation coverage, identifies information gaps, and automatically generates follow-up queries. The system typically conducts 3-4 search rounds, ensuring key dimensions (education, career, social, legal) have sufficient coverage.
Cross-Validation
Facts are verified across independent sources. Each fact is labeled as "corroborated" (2+ independent sources), "single-source," or "contested" (conflicting sources). Validation uses entity matching and timeline consistency checks.
Report Synthesis
Generates a structured investigation dossier containing background and education, work experience, social profiles, connections, timeline, relationship graph, evidence sources (E1-E5 classification), and trust score (0-100%).
Evidence Classification System
Seme uses a five-tier evidence classification system to evaluate the credibility of each information source. This system helps users quickly assess the reliability of each finding in the investigation results.
| Level | Name | Description |
|---|---|---|
| E1 | Primary Authoritative Sources | Official records, government documents, court decisions, institutional confirmations. Highest credibility. |
| E2 | Mainstream News Reports | Reports from established news media, corporate announcements, regulatory filings. |
| E3 | Social Media Posts | Content published on personal social media, forum posts, online reviews. |
| E4 | Anonymous Sources | Anonymous tips, unofficial channel information, unverified rumors. |
| E5 | Unverified Claims | Single-source claims that cannot be cross-validated. Lowest credibility. |
Trust Scoring
The trust score is a comprehensive 0-100% confidence metric calculated based on three weighted factors:
E1 sources weighted 1.0, E2 at 0.8, E3 at 0.5, E4 at 0.3, E5 at 0.1.
Proportion of corroborated facts to total facts. Facts with 2+ source corroboration weighted higher.
Coverage degree of key dimensions (education, career, social, legal).
Seme vs Traditional Methods
| Dimension | Traditional | Seme |
|---|---|---|
| Sources | 1-3 databases | 15-20 independent sources |
| Verification | Single source | Multi-source cross-validation |
| Evidence | No standard | E1-E5 五级体系 |
| Time | Days to weeks | 10-30 minutes |
| Output | Unstructured | Structured dossier + trust score |
| Auditability | Low | High (every fact sourced) |
Data Sources and Coverage
Seme's investigation engine integrates multiple data source categories to ensure coverage across multiple life and professional dimensions of the subject. Primary data source categories include: search engine results (aggregated via SearXNG from 70+ search providers), social media platforms (LinkedIn, Twitter/X, GitHub, Weibo, etc.), academic databases (Google Scholar, DBLP, ResearchGate), public records (company registrations, court documents, government announcements), news archives (major global news organizations), and domain registration information (WHOIS records).
Each data source has different update frequencies and coverage ranges. Search engines provide the broadest coverage but may contain outdated information. Social media provides real-time data but only reflects content users actively share. Public records provide authoritative official information but update more slowly. Seme's cross-validation mechanism compensates for individual source limitations by comparing data across sources, ensuring the final dossier information is both comprehensive and accurate.
Methodology Limitations
Despite Seme's carefully designed investigation methodology, there are inherent limitations users should be aware of: First, the system relies solely on publicly available information (OSINT principles) and cannot access private databases, paywalled content, or sources requiring authorization. Second, information timeliness depends on the update frequency of each data source, and some public records may have delays of weeks to months. Third, facial recognition accuracy is affected by image quality, lighting conditions, and facial occlusion, with real-world accuracy (95-98%) being lower than benchmark data (99.5%).
Additionally, the completeness of investigation results depends on the size of the subject's online footprint. Public figures with rich public information typically receive more comprehensive dossiers, while individuals who deliberately maintain a low profile may have limited results. The trust score reflects the credibility of currently available information but should not be treated as an absolute evaluation of the subject. Users should treat Seme's investigation results as one reference for decision-making, not the sole basis.
Platform Verified Data
User Feedback
"Seme reduced our background check time from 3 days to 20 minutes while covering more data sources. The E1-E5 evidence classification makes auditing simple."
"As a journalist, I need to quickly verify information sources. Seme's cross-validation and trust scores help me quickly assess information reliability."
Related Resources
- NATO OSINT Handbook — Official definition and methodology of open source intelligence
- NIST Identity Verification Standards — NIST identity verification framework
- Seme vs Traditional Methods Comparison — Detailed comparison across 12 dimensions
- Identity Investigation Glossary — 25+ OSINT and identity investigation term definitions