Deep Research
Multi-round AI investigation for comprehensive identity dossiers
Deep Research is a multi-round AI investigation process that systematically explores multiple angles of a subject's background. It follows a structured methodology: research planning, parallel search execution across 15-20 independent sources, gap analysis, cross-validation (2+ sources per fact), and report synthesis. A typical investigation involves 3-4 rounds, producing a structured dossier with education, career, social profiles, connections, timeline, evidence classification (E1-E5), and trust score. The investigation engine begins by generating a research plan that identifies 5-8 investigation angles (professional history, social presence, biographical data, education verification, association mapping, media coverage, legal records, and financial connections). Each angle is assigned to a dedicated search agent that queries multiple sources in parallel using both SearXNG meta-search (aggregating 70+ search sources) and OpenAI Web Search API. After each round, a gap analysis module evaluates coverage across all angles, identifies missing information, and generates targeted follow-up queries for the next round. Cross-validation is the core differentiator: every extracted fact is compared against independent sources, classified as corroborated (2+ independent sources), single-source (1 source only), or contested (conflicting sources). Corroborated facts receive high confidence weighting (5x for E1 government sources, 3x for E2 authoritative media, 2x for E3 professional platforms), while single-source facts are flagged for manual review. The final trust score is calculated as a weighted average: source quality (40%), cross-validation rate (35%), and information completeness across 8 dossier sections (25%). A score above 80% indicates high confidence with minimal gaps; 60-80% suggests moderate confidence with some unverified claims; below 60% indicates significant information gaps requiring additional investigation.
Technical Architecture and How It Works
Deep Research is built on modern AI infrastructure, leveraging deep learning, natural language processing, and distributed search technologies to deliver high-precision results. The system uses a microservices architecture, separating query parsing, data retrieval, result ranking, and dossier generation into independent service components, ensuring each stage can be independently scaled and optimized. The frontend uses Next.js 16 Server Components for server-side rendering, ensuring AI crawlers and search engines can fully index all content. The backend is deployed on Cloudflare Workers' global edge network, ensuring low-latency responses from any geographic location.
The data retrieval layer integrates multiple search providers, including the SearXNG meta-search engine (aggregating 70+ search sources), OpenAI Web Search API, and specialized academic and social media data sources. Each query is dispatched in parallel to multiple search agents, with results deduplicated, ranked, and cross-validated at the aggregation layer. The system uses vector embeddings (128-512 dimensions) for semantic similarity matching, combined with cosine similarity algorithms to ensure matching precision. For identity verification scenarios, the system additionally uses biometric feature comparison and timeline consistency checks.
The report generation stage uses GPT-4 for structured output synthesis, converting raw facts into investigation dossiers with evidence classification (E1-E5). Each fact is annotated with source URL, collection time, and verification status. The trust score is calculated based on three core factors: source quality (primary authoritative sources weighted highest), cross-validation rate (proportion of multi-source corroborated facts), and information completeness (coverage of key dimensions). The final dossier contains eight structured sections: background and education, work experience, social profiles, connections, timeline, relationship graph, evidence and sources, and trust score.
Workflow
Key Statistics
How It Works
Provide Input
Enter a person's name and any known background — employer, education, location, or any identifying information.
Research Planning
AI designs a multi-angle investigation plan covering professional history, social presence, biographical data, education, and associations.
Multi-Round Search
3-4 rounds of parallel searches across SearXNG, web search APIs, academic databases, and social platforms. Each round builds on previous findings.
Cross-Validation
Every fact is verified across independent sources. Corroborated facts (2+ sources) receive high confidence; single-source facts are flagged.
Dossier Delivery
Receive a structured dossier with 8 components: background, work history, social profiles, connections, timeline, evidence, relationship graph, and trust score.
Performance Comparison
Seme vs Traditional Methods
| Dimension | Traditional | Seme |
|---|---|---|
| Sources per Investigation | 1-3 databases | 15-20 independent sources |
| Verification Method | Single-source lookup | Cross-validation (2+ sources) |
| Evidence Classification | None | E1-E5 five-tier system |
| Investigation Rounds | 1 (static query) | 3-4 adaptive rounds |
| Time to Complete | Days to weeks | 10-30 minutes |
| Output Format | Unstructured notes | Structured dossier + trust score |
Use Cases
Executive Hiring
Comprehensive background investigation for C-level candidates — verify employment history, education, publications, and uncover potential red flags.
Investment Due Diligence
Investigate startup founders and key personnel before investment — verify claimed exits, check litigation history, validate credentials.
Compliance Screening
KYC/AML compliance screening for financial institutions — comprehensive background verification with source-cited evidence.
Investigative Journalism
Deep background research on subjects of investigative reports — verify claims, map connections, build evidence-based profiles.
Security and Compliance
Seme is designed with data security and privacy compliance requirements in mind. All data transfers use TLS 1.3 encryption, and static data uses AES-256 encryption. The system only collects information from publicly available sources (OSINT principles), never accessing private databases or content requiring authorization. Investigation results are stored in the user's workspace following the principle of least privilege. The platform complies with GDPR data processing requirements, and users can export or delete their investigation data at any time.
For enterprise customers, Seme supports single sign-on (SSO), audit logs, role-based access control, and data residency options. API access uses OAuth 2.0 authentication, and all API calls have rate limiting and usage auditing. The system undergoes regular security audits and penetration testing to ensure compliance with industry security standards.
Industry Applications
Financial Services
Banks and investment firms use Seme for KYC/AML compliance checks, anti-fraud investigations, and counterparty due diligence. The system can complete background checks in 10-30 minutes that traditionally take days, significantly reducing compliance costs and risk exposure time.
Legal and Investigation
Law firms and investigation agencies use Seme for witness background checks, asset tracing, and relationship analysis. Structured dossier outputs can be directly used in legal documents, with all evidence annotated with source and credibility levels.
Human Resources
HR teams use Seme to verify candidate education, work history, and professional credentials, detecting resume fraud and misrepresentation. The semantic search function helps recruiters discover passive candidates matching specific criteria.
Why Choose Seme
Seme has helped security teams, journalists, and HR professionals complete thousands of identity investigations. Compared to traditional background checks, Seme reduces investigation time from days to minutes while providing more comprehensive data coverage and auditable evidence chains.
Related Resources
Learn more about identity investigation methodology and OSINT techniques:
Frequently Asked Questions
How many sources does deep research use?▾
A typical deep research investigation queries 15-20 independent sources across search engines, social media platforms, academic databases, public records, and news archives. The system performs 3-4 rounds of searches, with each round building on findings from the previous one.
What is the trust score?▾
The trust score is a 0-100% confidence percentage based on source quality and corroboration. E1 sources (government records) carry the highest weight (5x), while E5 sources (unverified claims) carry the lowest (0.5x). A score above 80% indicates high confidence; 60-80% is moderate; below 60% has significant gaps.
Can I customize the investigation scope?▾
Yes. You can provide additional context to guide the investigation — specific aspects to focus on, known organizations to investigate, or particular time periods to examine. The AI will incorporate your guidance into its research plan while also exploring standard investigation angles.
What is the E1-E5 evidence classification system?▾
The E1-E5 system classifies evidence by source reliability: E1 (Government/Official records) — highest reliability, 5x weight; E2 (Authoritative media, established institutions) — high reliability, 3x weight; E3 (Professional platforms like LinkedIn, academic databases) — moderate reliability, 2x weight; E4 (Social media, user-generated content) — lower reliability, 1x weight; E5 (Unverified claims, anonymous sources) — lowest reliability, 0.5x weight. The trust score calculation uses these weights to prioritize high-quality evidence.
How does deep research handle conflicting information?▾
When the system encounters conflicting facts from different sources, it flags them as "contested" in the dossier. Each conflicting claim is presented with its source and evidence level, allowing you to evaluate which version is more credible. The system prioritizes higher-tier sources (E1-E2) over lower-tier ones (E4-E5) and notes the number of sources supporting each version. For critical contradictions (e.g., conflicting employment dates), the system generates a follow-up query specifically targeting resolution of that conflict.