Last verified: October 9, 2026.
Unstructured data is information that does not conform to a predefined data model or schema. Approximately 80-90% of enterprise data is unstructured, and the proportion is growing as organizations generate more text, image, video, audio, and sensor data. The eight main categories of unstructured data are text documents, images, video, audio, social media, web pages, log files, and IoT sensor data. Storing unstructured data requires object storage (S3, GCS, ADLS); processing it requires specialized tools (Spark, Databricks, Snowflake, vector databases, multimodal LLMs). This walkthrough reads unstructured data the way a data team encountering it reads the format: the eight categories, the file formats within each, the storage and processing tools, the comparison against structured and semi-structured data, and the 2026 trends (multimodal LLMs, vector search, lakehouse architectures) that are reshaping how organizations manage unstructured data.
The 8 Categories of Unstructured Data
Unstructured data spans eight main categories. Each category has its own file formats, storage requirements, and processing tools:
| Category | Examples | Typical size | Storage cost per GB/month |
|---|---|---|---|
| Text documents | PDF, Word, Google Docs, Notion, emails, chat logs | 100KB-50MB per document | $0.02-$0.05 |
| Images | JPEG, PNG, TIFF, DICOM (medical), satellite imagery | 2MB-100MB per image | $0.02-$0.05 |
| Video | MP4, MOV, security footage, drone footage, training videos | 100MB-10GB per minute | $0.02-$0.10 |
| Audio | MP3, WAV, FLAC, call recordings, voice notes, podcasts | 1MB-50MB per minute | $0.02-$0.05 |
| Social media | Tweets, LinkedIn posts, Instagram captions, TikTok, Reddit threads | 1KB-10MB per post | $0.05-$0.20 |
| Web pages | HTML, JSON API responses, scraped pages, JSON-LD structured data | 50KB-5MB per page | $0.02-$0.10 |
| Log files | Server logs, application logs, audit logs, security event logs, clickstream | 1GB-1TB per day per service | $0.01-$0.05 |
| IoT sensor data | Equipment telemetry, smart home events, vehicle telematics, environmental sensors | 10GB-100GB per day per fleet | $0.01-$0.05 |
Log files and IoT sensor data dominate by volume: a single large web service can produce 1-10 TB of logs per day, and a manufacturing plant can produce 100+ GB of sensor data per day. Text documents and emails dominate by count: an enterprise with 1,000 employees produces roughly 5-10 million emails per year. Images and video dominate by storage cost: medical imaging and security footage are the largest per-byte storage categories.
Structured vs Semi-Structured vs Unstructured
The three data categories are distinguished by the presence and consistency of the schema:
| Type | Schema | Examples | Storage | Query |
|---|---|---|---|---|
| Structured | Predefined schema (tables, columns, types) | Customer records, financial transactions, sensor measurements | SQL databases (Postgres, MySQL, Snowflake, BigQuery, Redshift) | SQL |
| Semi-structured | Self-describing schema (tags, fields) | JSON, XML, YAML, Parquet with metadata, Avro | Document stores (MongoDB, DynamoDB, CosmosDB); columnar (Parquet, ORC) | Document query languages, SQL with JSON support |
| Unstructured | No predefined schema; format implies structure | PDF, JPEG, MP4, MP3, log files, raw text | Object storage (S3, GCS, ADLS); data lakes; vector databases (Pinecone, Weaviate, Milvus) | Custom parsing, AI/ML, vector search, multimodal LLMs |
The boundary between semi-structured and unstructured is fuzzy. JSON and XML are technically self-describing (semi-structured), but the nested structure is irregular and requires parsing. A scanned PDF is fully unstructured (no machine-readable text); a digitally-generated PDF with text layers is semi-structured (text is extractable but the format is irregular). The right framing is that unstructured data requires parsing or AI to extract structure; semi-structured data has the structure embedded but is irregular.
The lakehouse architecture (Databricks, Snowflake, BigQuery) blurs the boundary further by adding unstructured-data processing capabilities to the data warehouse. Snowflake's unstructured-data support, Databricks' Delta Lake with multimodal processing, and BigQuery's object-table integration are examples of the convergence.
Storage: The Object Storage Layer
Unstructured data lives in object storage, not in traditional databases:
| Provider | Service | Storage cost per GB/month | Egress cost per GB | Best for |
|---|---|---|---|---|
| AWS | S3 Standard | $0.023 | $0.09 | General-purpose, AWS ecosystem |
| AWS | S3 Infrequent Access | $0.0125 | $0.09 | Disaster recovery, archives accessed monthly |
| AWS | S3 Glacier Deep Archive | $0.00099 | $0.09 | Long-term compliance archives |
| Google Cloud | Cloud Storage Standard | $0.020 | $0.12 | General-purpose, GCP ecosystem |
| Google Cloud | Cloud Storage Nearline | $0.010 | $0.12 | Infrequent access |
| Google Cloud | Cloud Storage Coldline | $0.004 | $0.12 | Cold storage, accessed quarterly |
| Azure | Blob Storage Hot | $0.0184 | $0.087 | General-purpose, Azure ecosystem |
| Azure | Blob Storage Cool | $0.010 | $0.087 | Infrequent access |
| Azure | Blob Storage Archive | $0.00099 | $0.087 | Long-term compliance archives |
| Cloudflare | R2 (no egress fees) | $0.015 | $0.00 | Egress-heavy workloads (CDN, large datasets) |
| Backblaze | B2 | $0.006 | $0.01 | Cost-sensitive archives |
The S3 pricing is the industry benchmark. S3 Standard at $0.023/GB/month is the right pick for actively-used data; S3 IA at $0.0125 is the right pick for monthly-accessed archives; S3 Glacier at $0.00099 is the right pick for compliance archives. The egress cost ($0.09/GB) is the right pick for applications that read data internally (the same AWS region), but the egress cost is a major expense for cross-region or cross-cloud data movement.
Cloudflare R2 with zero egress fees is the right pick for egress-heavy workloads: serving media files through a CDN, large-dataset distribution, or backup-and-restore scenarios. R2's storage is slightly cheaper than S3 Standard; the egress savings are the major benefit.
Backblaze B2 at $0.006/GB/month is the right pick for cost-sensitive archives where the egress cost is minimal. B2 is roughly 1/4 the cost of S3 Standard, but the integration with the rest of the data stack is weaker; B2 is typically the right pick for backups, not for primary data.
Processing: The Tooling Layer
Unstructured data processing tools in 2026, by category:
| Tool | What it does | Best for |
|---|---|---|
| Apache Spark | Distributed batch processing for large datasets | ETL, log analysis, image processing, large-scale data transformations |
| Databricks Lakehouse Platform | Spark + Delta Lake + MLflow + Mosaic AI | Unified analytics + ML on unstructured + structured data |
| Snowflake (unstructured support) | Object storage + Cortex AI for document/AI processing | Unified structured + unstructured queries from one platform |
| BigQuery + Vertex AI | GCP-native data warehouse + multimodal AI | GCP shops, multimodal data (image, video, text) |
| AWS Glue + Bedrock + Rekognition | AWS-native ETL + foundation models + image/video analysis | AWS shops, computer-vision workloads |
| Vector databases (Pinecone, Weaviate, Milvus, Qdrant) | Embeddings storage + semantic search | RAG pipelines, semantic search, recommendation systems |
| OpenSearch / Elasticsearch | Full-text search + log analytics + vector search | Search engines, log analytics, hybrid keyword+vector search |
| Multimodal LLMs (GPT-Realtime-2, Claude 4.5, Gemini 2.5, Grok 4.7) | Image / video / audio understanding + reasoning | Document extraction, image classification, video understanding, audio transcription |
The 2026 trend is the lakehouse convergence: data warehouses (Snowflake, BigQuery, Databricks SQL) are absorbing unstructured-data processing capabilities. Snowflake's Cortex AI for document and image processing, BigQuery's object-table integration with Vertex AI, and Databricks' Unity Catalog with multimodal AI are all examples of the convergence. The right pick for new projects in 2026 is the lakehouse platform that handles both structured and unstructured data in one place.
Vector databases are the right pick for RAG (retrieval-augmented generation) and semantic search workloads. The vector database stores embeddings of the unstructured data (text, image, audio chunks), and the LLM retrieves the relevant chunks for the prompt. The architecture is the foundation for 2026's AI applications on enterprise data.
Multimodal LLMs are the right pick for direct analysis of unstructured data. GPT-Realtime-2, Claude 4.5 Opus, Gemini 2.5 Pro, and Grok 4.7 can all process images, video, and audio directly; the right pattern for new AI applications is to send the unstructured data to the multimodal LLM and use the response as the analysis output.
The 2026 Trend: Multimodal AI on Unstructured Data
Three trends are reshaping how organizations manage unstructured data in 2026:
- Multimodal LLMs as the analysis layer. GPT-Realtime-2, Claude 4.5 Opus, Gemini 2.5 Pro, and Grok 4.7 can all process text, image, video, and audio directly. The pattern is to send the unstructured data to the multimodal LLM and use the response as the structured output. The custom-parsing-and-AI pipeline that dominated 2020-2023 is being replaced by end-to-end multimodal LLMs.
- Vector search as the retrieval layer. Embeddings of unstructured data (text chunks, image embeddings, audio embeddings) are stored in vector databases; the LLM retrieves the relevant chunks for the prompt. The architecture is the foundation for RAG, semantic search, and recommendation systems on enterprise data.
- Lakehouse architecture as the data foundation. Snowflake, BigQuery, Databricks are absorbing unstructured-data processing. The pattern is to store the unstructured data in object storage, register it in the lakehouse catalog, and query it with SQL or AI tools. The traditional "data warehouse + separate data lake + separate AI platform" architecture is converging.
The combined effect is that unstructured data is no longer a separate problem. The 2026 data platform handles structured, semi-structured, and unstructured data through one interface; the analysis layer is multimodal LLMs; the retrieval layer is vector search; the storage layer is the lakehouse. Organizations that have not yet adopted the converged architecture are spending 2-5x more on data infrastructure than the ones that have.
What to Budget For
Three cost paths settle the unstructured-data decision.
Small projects (under 1 TB): budget $50-$200/month for S3 Standard storage + $100-$500/month for processing (Spark on a small cluster, or Snowflake/BigQuery on-demand). Total data infrastructure: $200-$700/month.
Mid-market (1-100 TB): budget $1,000-$5,000/month for object storage + $5,000-$20,000/month for processing. Total data infrastructure: $10,000-$30,000/month. The lakehouse architecture (Databricks, Snowflake, or BigQuery) is the right pick for this scale.
Enterprise (100+ TB): budget $5,000-$50,000/month for object storage + $50,000-$500,000/month for processing. Total data infrastructure: $100,000-$1,000,000/month. The right pick is a multi-region object storage + a lakehouse platform + a vector database + a multimodal AI layer.
Add $50,000-$500,000/year for engineering team to build the pipelines, the vector search infrastructure, and the multimodal AI integrations. The cost of the data infrastructure is the smaller share of the total cost; the engineering team is the larger share.
What to Skip
Three cost traps to skip.
Storing unstructured data in databases. The cost per GB of object storage ($0.02-$0.05) is 10-50x cheaper than the cost per GB of database storage. Object storage is the right place for unstructured data; databases are the right place for structured data.
Custom image / video / audio processing pipelines in 2026. Multimodal LLMs handle the analysis; building a custom pipeline is engineering investment that the multimodal LLM commoditizes. The right pattern is to send the unstructured data to a multimodal LLM and use the response as the structured output.
Vendor-locked unstructured data platforms. The lakehouse convergence means the data warehouse, the data lake, and the AI platform are converging. Picking a vendor-locked unstructured data platform creates switching cost without a corresponding benefit. The right pick is the lakehouse platform that handles all three data types.
Read next
DBT Data Build Tool: Data Engineering Explained is the structured-data transformation framework that handles the structured half of the lakehouse architecture, and What Is Data Engineering 2026: Role, Tools, Salary is the broader data-engineering context that places unstructured data within the modern data stack.






