15-20 sources per investigation

Identity Investigation Glossary

Definitions of terms and concepts used in identity investigation, open source intelligence (OSINT), biometrics, and due diligence.

This glossary covers core concepts in the identity investigation field, including OSINT methodology, biometric technologies, evidence classification systems, trust scoring mechanisms, and compliance frameworks. Each term provides bilingual definitions and is linked to specific Seme platform features and use cases. Terms are organized by category for easy reference and understanding.

Identity investigation is a rapidly evolving field that combines multiple disciplines including Open Source Intelligence (OSINT), artificial intelligence, biometrics, and data analytics. Understanding these professional terms is essential for effectively using investigation tools, interpreting results, and meeting compliance requirements. The Seme platform integrates these concepts into a unified workflow, enabling non-specialist users to conduct professional identity investigations. Whether you are a security analyst, HR professional, journalist, or investor, mastering these terms will help you better leverage AI investigation tools.

Identity investigation knowledge framework overview
Intelligence (2)Technology (8)Methodology (6)Security (2)Concepts (3)Use Cases (3)Data Sources (1)
Intelligence2Technology8Methodology6Security2Concepts3Use Cases3Data Sources1

Intelligence

OSINT (Open Source Intelligence)

OSINT (Open Source Intelligence) is the practice of collecting and analyzing information from publicly available sources to produce actionable intelligence. According to NATO guidelines, OSINT sources include social media platforms, public records databases, news archives, academic publications, government filings, and domain registration records. In identity investigation, OSINT combines data from LinkedIn profiles, company filings, academic publications, and social media activity to construct comprehensive background profiles. Modern OSINT workflows use AI to cross-reference findings across 15-20 independent sources, achieving verification confidence levels from E1 (multiple corroborating primary sources) to E5 (single unverified source). The U.S. Department of Defense defines OSINT as "information that has been obtained from publicly available sources and has been validated through intelligence tradecraft."

Social Media Intelligence (SOCMINT)

Social Media Intelligence (SOCMINT) is a specialized subset of OSINT focused on collecting and analyzing information from social media platforms. SOCMINT encompasses monitoring public posts, analyzing network connections, tracking behavioral patterns, extracting structured data from social profiles, and assessing sentiment and influence. Key platforms analyzed include Twitter/X, LinkedIn, Facebook, Instagram, Reddit, Weibo, and TikTok. SOCMINT differs from general social media monitoring in its analytical depth — it goes beyond tracking mentions to map relationship networks, identify influence patterns, detect behavioral anomalies, and construct timeline narratives from social activity.

Technology

Face Search

Face Search is a biometric identification technique that uses AI-powered deep learning to match a facial photograph against databases of known faces. Modern face search systems convert facial images into high-dimensional embedding vectors (typically 128-512 dimensions) using convolutional neural networks (CNNs) such as FaceNet, ArcFace, or DeepFace. These embeddings capture unique facial features — the distance between eyes, nose shape, jawline contour, and skin texture — allowing matches even with variations in lighting, angle, expression, and aging. State-of-the-art systems achieve 99.5%+ accuracy on benchmark datasets like LFW (Labeled Faces in the Wild). In identity investigation, face search bridges the gap between having a photograph and knowing who the person is.

Semantic Search

Semantic Search is a search methodology that understands the intent and contextual meaning behind a query rather than just matching keywords. Unlike traditional keyword search, which relies on exact term matching and TF-IDF scoring, semantic search uses natural language understanding (NLU) models — typically transformer-based architectures like BERT or GPT — to interpret query intent and match it against semantically relevant results. In identity investigation, semantic search interprets natural language criteria such as "all researchers who worked at Tencent AI Lab between 2018 and 2022" to find people matching specific descriptions, even when no exact keyword match exists in the data.

Knowledge Graph

A Knowledge Graph is a network representation of entities (people, organizations, locations) and the relationships between them, stored as a graph of nodes and edges. In identity investigation, knowledge graphs visualize connections discovered during research, helping analysts understand complex networks of associations. Each node represents an entity with attributes (name, type, properties), and each edge represents a relationship (works_at, knows, published_with, located_in). Knowledge graphs enable graph-based queries such as "find all people who worked at both Company A and Company B" or "identify the shortest connection path between Person X and Person Y." Modern investigation platforms use knowledge graphs with thousands of nodes to map organizational structures, influence networks, and entity relationships.

Entity Resolution

Entity Resolution is the process of determining whether different records or data points refer to the same real-world entity. In identity investigation, entity resolution links profiles across platforms — matching a Twitter handle to a LinkedIn profile, connecting an email address to a real name, or associating a phone number with multiple online accounts. This is one of the most challenging aspects of identity investigation because people use different names, handles, and identifiers across platforms. Modern entity resolution uses probabilistic matching algorithms that consider multiple signals: exact matches (same email, phone), fuzzy matches (similar names, partial matches), behavioral signals (writing style, activity patterns), and contextual signals (shared connections, overlapping timelines).

Data Enrichment

Data Enrichment is the process of enhancing existing records with additional information from external sources to create a more complete and accurate profile. In identity investigation, starting with minimal input — a name, a photo, or an email address — enrichment progressively adds social profiles, employment history, education records, public records, contact information, and other contextual data. The enrichment process follows a funnel pattern: starting broad (searching across all available sources) and narrowing based on confirmed matches. Each enrichment step increases the confidence score of the overall profile. Modern enrichment platforms can take a single identifier and expand it into a comprehensive 50+ field profile in minutes.

Biometric Identification

Biometric Identification uses unique physical or behavioral characteristics to identify and verify individuals. The primary biometric modalities in identity investigation are: facial recognition (most common, using CNN-based face embeddings), fingerprint analysis (ridge pattern matching), voice recognition (vocal feature extraction), iris scanning (iridial pattern analysis), and gait analysis (walking pattern recognition). Facial recognition dominates identity investigation due to the abundance of facial photographs available from social media, surveillance systems, and public records. Modern facial recognition systems achieve 99.5%+ accuracy on controlled datasets and 95-98% accuracy on real-world images, with performance varying by lighting conditions, image quality, and demographic factors.

Alias Detection

Alias Detection is the process of identifying alternative names, handles, or identities used by the same individual across different platforms and contexts. People use aliases for various reasons — privacy, professional separation (personal vs. work accounts), gaming personas, or deliberate concealment. Alias detection relies on five signal categories: (1) Shared Identifiers — matching email addresses, phone numbers, or recovery emails across accounts; (2) Username Patterns — detecting similar naming conventions (jdoe, john.doe, johndoe82); (3) Behavioral Signals — matching writing style, vocabulary, and topic interests; (4) Visual Signals — matching profile photos through perceptual hashing or facial recognition; (5) Network Overlap — detecting shared connections or community membership. Effective alias detection can uncover 3-10 additional online identities per subject.

Open Source Intelligence Tools

Open Source Intelligence Tools are software applications and platforms designed to facilitate OSINT collection, processing, and analysis. The OSINT tool ecosystem includes: search aggregators (SearXNG, which queries multiple engines simultaneously), social media analysis tools (social media scrapers, profile analyzers), reverse image search engines (Google Images, TinEye, Yandex), network analysis frameworks (Maltego, Gephi), public record databases (PACER, EDGAR, Companies House), data visualization platforms, and specialized investigation platforms like Seme. Modern OSINT tools increasingly incorporate AI for automated entity extraction, relationship mapping, and pattern recognition, reducing the manual effort required for comprehensive investigations.

Methodology

Deep Research

Deep Research is a multi-round AI investigation process that systematically explores multiple angles of a subject's background to produce a comprehensive identity dossier. Unlike single-query searches, deep research follows a structured methodology: research planning (designing investigation angles — professional, social, biographical, educational), parallel search execution (running multiple queries simultaneously across search engines and databases), gap analysis (evaluating coverage and generating follow-up queries), cross-validation (verifying facts across independent sources), and report synthesis (generating a structured dossier). A typical deep research investigation involves 3-4 rounds, each building on findings from the previous round, covering 15-20 independent sources per subject.

Investigation Dossier

An Investigation Dossier is a structured compilation of all findings from an identity investigation, designed to present verified facts in a standardized, auditable format. A complete dossier contains eight core components: (1) Background and Education — academic history, degrees, institutions, research focus; (2) Work Experience — career timeline, roles, companies, responsibilities, achievements; (3) Social Media — verified social profiles with confidence scores; (4) Connections and Associations — related persons, organizations, affiliations; (5) Timeline — key life events chronologically ordered with source citations; (6) Relationship Graph — visual network of entity connections; (7) Evidence and Sources — all source URLs with E1-E5 classification; (8) Trust Score — overall confidence percentage based on source quality and cross-validation.

Cross-Validation

Cross-Validation in identity investigation is the process of verifying facts by confirming them across multiple independent sources. A fact supported by two or more unrelated sources is considered corroborated and receives a high confidence classification. Single-source facts are flagged with lower confidence and may trigger additional verification searches. Cross-validation follows three principles: (1) Independence — sources must not share a common origin (e.g., a LinkedIn post and a company press release that references that post count as one source, not two); (2) Consistency — corroborating sources must agree on key facts; (3) Recency — more recent sources receive higher weight for time-sensitive facts. The cross-validation engine assigns confidence levels: High (3+ independent sources), Medium (2 sources), Low (1 source), Contested (conflicting sources).

Evidence Classification (E1-E5)

Evidence Classification is a five-tier system for categorizing the reliability and quality of evidence sources in identity investigation. E1 (Primary Sources) includes government records, official filings, court documents, and verified institutional records — these carry the highest evidentiary weight. E2 (Verified Secondary Sources) includes reputable news articles, published academic papers, and verified corporate announcements. E3 (Social & User-Generated Content) includes verified social media posts, professional profiles, and public forum contributions. E4 (Unverified Sources) includes anonymous posts, unverified claims, and single-source allegations. E5 (Lowest Confidence) includes rumors, speculation, and content from sources with known reliability issues. This classification system enables investigators to assess the overall confidence of their findings and identify areas requiring additional verification.

Network Analysis

Network Analysis is the study of relationships and connections between entities to understand social structures, information flow, and influence patterns. In identity investigation, network analysis maps professional relationships, organizational affiliations, and social connections using graph theory metrics. Key metrics include: degree centrality (number of direct connections), betweenness centrality (how often an entity lies on the shortest path between others), clustering coefficient (how interconnected an entity's contacts are), and eigenvector centrality (connection quality based on the importance of connected entities). Network analysis reveals hidden structures that are not visible from examining individual profiles — for example, a person with moderate direct connections but high betweenness centrality may be a critical information broker.

Timeline Analysis

Timeline Analysis is the process of organizing discovered facts chronologically to create a coherent narrative of a subject's life events. In identity investigation, timeline analysis transforms scattered data points from multiple sources into a structured chronological sequence, helping investigators identify gaps, inconsistencies, and patterns that may require further investigation. A well-constructed timeline includes dated events (employment start/end, education enrollment/graduation, publication dates, travel events), source citations for each event, and confidence indicators. Timeline analysis is particularly valuable for verifying alibis, detecting fabrication (events that don't fit the timeline), and identifying periods of unexplained activity or inactivity.

Security

Identity Verification

Identity Verification is the process of confirming that a person is who they claim to be by cross-referencing multiple independent data points. In digital contexts, verification combines biometric data (facial recognition, fingerprint), documentation (government IDs, academic credentials), behavioral patterns (writing style, activity patterns), and social signals (professional connections, organizational affiliations) to establish confidence in an identity. Modern identity verification systems use a tiered approach: Level 1 checks basic document validity, Level 2 cross-references with external databases, and Level 3 performs biometric matching. The confidence level is expressed as a percentage based on the number and quality of corroborating sources.

Data Breach

A Data Breach is an incident where protected or confidential data is accessed, disclosed, or stolen by unauthorized parties. Common breach types include hacking, insider threats, physical theft, and accidental exposure. Breached data — often found on dark web marketplaces, paste sites, and breach notification services like HaveIBeenPwned — can be a valuable source of investigative information, though its use raises significant ethical and legal considerations. In identity investigation, breach data may reveal aliases, email addresses, password patterns, account associations, and previously unknown online identities. The ethical use of breach data requires careful consideration of privacy laws, data protection regulations, and the distinction between publicly available breach indexes and raw leaked data.

Concepts

Digital Footprint

A Digital Footprint is the comprehensive trail of data and information left behind by a person's online activities. It encompasses both active footprints (social media posts, blog articles, forum comments, profile registrations) and passive footprints (cookies, IP logs, location data, device fingerprints). In identity investigation, analyzing a person's digital footprint reveals their professional network, interests, location history, behavioral patterns, and associations. A typical adult's digital footprint spans 15-25 distinct platforms and services, generating hundreds of data points that can be aggregated into a coherent identity profile. Digital footprints are classified by permanence: ephemeral (stories, live streams), semi-permanent (social media posts), and permanent (public records, archived content).

Trust Score

A Trust Score is a confidence percentage assigned to an investigation dossier based on the quality, quantity, and corroboration of its sources. The score ranges from 0% (no verified information) to 100% (all facts fully corroborated by multiple high-quality sources). The calculation considers: (1) Source Quality — weighted by evidence classification tier (E1=5x, E2=3x, E3=2x, E4=1x, E5=0.5x); (2) Corroboration Rate — percentage of facts supported by 2+ independent sources; (3) Coverage Completeness — how many dossier sections contain verified data; (4) Consistency — absence of conflicting facts. Scores above 80% indicate high confidence suitable for decision-making; 60-80% indicates moderate confidence with some gaps; below 60% indicates significant verification gaps.

Person of Interest

A Person of Interest (POI) is an individual who has come to the attention of investigators or analysts for a specific reason, but who may not necessarily be suspected of wrongdoing. In identity investigation, a POI is typically someone whose background needs to be verified, whose connections need to be mapped, or whose activities need to be monitored as part of a larger investigation. POI designation is the starting point for focused investigation — it triggers a systematic process of data collection, analysis, and verification. The term is used across domains: law enforcement (persons connected to a case), corporate security (individuals flagged for access review), journalism (sources and subjects of investigation), and due diligence (persons requiring background verification).

Use Cases

Due Diligence

Due Diligence is a comprehensive appraisal of a person or organization undertaken before entering into a business relationship or transaction. Modern due diligence combines traditional background checks with OSINT techniques to verify credentials, assess risks, and uncover potential red flags. In the context of identity investigation, personal due diligence covers: employment verification, education confirmation, criminal record checks, litigation history, financial standing, regulatory sanctions, media reputation analysis, and association mapping. Due diligence depth varies by use case: basic screening (automated database checks), standard investigation (OSINT + database), and enhanced investigation (multi-round deep research with cross-validation). Regulatory frameworks like KYC/AML mandate due diligence for financial services, while voluntary due diligence is common in M&A, hiring, and partnership decisions.

Background Check

A Background Check is a systematic review of a person's commercial, criminal, financial, and personal records to verify their identity and assess their trustworthiness. AI-powered background checks go beyond traditional database lookups by incorporating OSINT data, social media analysis, and cross-referencing multiple sources for comprehensive verification. Traditional background checks rely on database queries (criminal records, credit reports, employment verification services) and typically cover 5-10 data points. AI-enhanced background checks add digital footprint analysis, social media screening, network analysis, and cross-validation across 15-20+ sources, producing a more comprehensive and nuanced assessment.

Competitive Intelligence

Competitive Intelligence (CI) is the practice of gathering and analyzing information about competitors, market trends, and industry dynamics to support strategic decision-making. Identity investigation tools support CI by mapping talent movements, organizational changes, and key personnel networks across companies. Unlike industrial espionage, competitive intelligence relies exclusively on legal and ethical information sources — public records, published reports, social media, patent filings, and academic publications. In the context of identity investigation, CI focuses on understanding the people behind competitive dynamics: who leads competitor teams, what expertise they have, how talent flows between companies, and what networks connect industry leaders.

Data Sources

Platform Data

25+
Glossary Terms
5
Evidence Levels
78%
Avg Trust Score
15-20
Sources/Investigation

Related Resources

Frequently Asked Questions

What is OSINT?

OSINT (Open Source Intelligence) is the practice of collecting and analyzing information from publicly available sources to produce actionable intelligence. Sources include social media, public records, news archives, academic publications, and government documents.

What is a trust score?

A trust score is a 0-100% confidence metric based on source quality and cross-validation results. Facts corroborated by multiple independent sources score higher.

What are evidence levels E1-E5?

E1-E5 is a five-tier system: E1 primary authoritative sources, E2 news articles, E3 social media, E4 anonymous sources, E5 unverified claims.

Get Started with Seme

Put these concepts into practice. Try the AI-powered identity investigation platform for free.