15-20 sources per investigation

Investigation Methodology

By Seme Research Team · Updated May 22, 2026

Seme investigation methodology workflow diagram

Seme's identity investigation methodology is based on Open Source Intelligence (OSINT) principles, combining AI-driven automated search with structured analysis. This methodology ensures every investigation undergoes systematic multi-source verification, producing auditable and reproducible results. Unlike traditional background checks that rely on single databases, Seme searches across 15-20 independent sources in parallel, confirms facts through cross-validation, and quantifies the credibility of each finding through a five-tier evidence classification system (E1-E5).

Investigation Workflow

InputPhoto/Name/CriteriaResearchPlanning5-8 directionsMulti-RoundSearch15-20 sourcesGapAnalysis3-4 roundsCross-ValidationCorroborationReportDossier+Score

Five-Step Investigation Process

1

Research Planning

AI analyzes input information and designs a multi-angle investigation strategy covering professional, social, academic, and public records dimensions. The system automatically generates 5-8 investigation directions, each with specific search queries and expected data sources.

2

Multi-Round Search

Parallel execution of SearXNG meta-search engine and OpenAI Web Search API queries, collecting information from 15-20 independent sources. Each round extracts structured facts including names, positions, education, social profiles, and affiliated organizations.

3

Gap Analysis

AI evaluates current investigation coverage, identifies information gaps, and automatically generates follow-up queries. The system typically conducts 3-4 search rounds, ensuring key dimensions (education, career, social, legal) have sufficient coverage.

4

Cross-Validation

Facts are verified across independent sources. Each fact is labeled as "corroborated" (2+ independent sources), "single-source," or "contested" (conflicting sources). Validation uses entity matching and timeline consistency checks.

5

Report Synthesis

Generates a structured investigation dossier containing background and education, work experience, social profiles, connections, timeline, relationship graph, evidence sources (E1-E5 classification), and trust score (0-100%).

Evidence Classification System

Seme uses a five-tier evidence classification system to evaluate the credibility of each information source. This system helps users quickly assess the reliability of each finding in the investigation results.

LevelNameDescription
E1Primary Authoritative SourcesOfficial records, government documents, court decisions, institutional confirmations. Highest credibility.
E2Mainstream News ReportsReports from established news media, corporate announcements, regulatory filings.
E3Social Media PostsContent published on personal social media, forum posts, online reviews.
E4Anonymous SourcesAnonymous tips, unofficial channel information, unverified rumors.
E5Unverified ClaimsSingle-source claims that cannot be cross-validated. Lowest credibility.

Trust Scoring

The trust score is a comprehensive 0-100% confidence metric calculated based on three weighted factors:

Source Quality40%Cross-Validation Rate35%Info Completeness25%
Source Quality
Weight: 40%

E1 sources weighted 1.0, E2 at 0.8, E3 at 0.5, E4 at 0.3, E5 at 0.1.

Cross-Validation Rate
Weight: 35%

Proportion of corroborated facts to total facts. Facts with 2+ source corroboration weighted higher.

Information Completeness
Weight: 25%

Coverage degree of key dimensions (education, career, social, legal).

Seme vs Traditional Methods

DimensionTraditionalSeme
Sources1-3 databases15-20 independent sources
VerificationSingle sourceMulti-source cross-validation
EvidenceNo standardE1-E5 五级体系
TimeDays to weeks10-30 minutes
OutputUnstructuredStructured dossier + trust score
AuditabilityLowHigh (every fact sourced)

Data Sources and Coverage

Seme's investigation engine integrates multiple data source categories to ensure coverage across multiple life and professional dimensions of the subject. Primary data source categories include: search engine results (aggregated via SearXNG from 70+ search providers), social media platforms (LinkedIn, Twitter/X, GitHub, Weibo, etc.), academic databases (Google Scholar, DBLP, ResearchGate), public records (company registrations, court documents, government announcements), news archives (major global news organizations), and domain registration information (WHOIS records).

Each data source has different update frequencies and coverage ranges. Search engines provide the broadest coverage but may contain outdated information. Social media provides real-time data but only reflects content users actively share. Public records provide authoritative official information but update more slowly. Seme's cross-validation mechanism compensates for individual source limitations by comparing data across sources, ensuring the final dossier information is both comprehensive and accurate.

Methodology Limitations

Despite Seme's carefully designed investigation methodology, there are inherent limitations users should be aware of: First, the system relies solely on publicly available information (OSINT principles) and cannot access private databases, paywalled content, or sources requiring authorization. Second, information timeliness depends on the update frequency of each data source, and some public records may have delays of weeks to months. Third, facial recognition accuracy is affected by image quality, lighting conditions, and facial occlusion, with real-world accuracy (95-98%) being lower than benchmark data (99.5%).

Additionally, the completeness of investigation results depends on the size of the subject's online footprint. Public figures with rich public information typically receive more comprehensive dossiers, while individuals who deliberately maintain a low profile may have limited results. The trust score reflects the credibility of currently available information but should not be treated as an absolute evaluation of the subject. Users should treat Seme's investigation results as one reference for decision-making, not the sole basis.

Platform Verified Data

15-20
Independent Sources/Investigation
78%
Average Trust Score
3-4
Investigation Rounds
10-30 min
Completion Time

User Feedback

"Seme reduced our background check time from 3 days to 20 minutes while covering more data sources. The E1-E5 evidence classification makes auditing simple."

S
Security Team Lead
Tech Company

"As a journalist, I need to quickly verify information sources. Seme's cross-validation and trust scores help me quickly assess information reliability."

I
Investigative Journalist
Media Organization

Related Resources