Dataset Python
Orders Dataset
A synthetic order dataset for practicing filters, aggregations, missing values, duplicate detection and multi-table pandas analysis.
- Type
- Dataset
- Level
- Beginner
- Updated
In short
Two hundred fifty order rows across several months, with repeated customers, status variation, nullable values, duplicates and unmatched keys.
Downloads and links
Who it is for
- Readers practicing transactional data analysis
- Data engineers learning validation before aggregation
What it helps you do
- Filter orders using multiple business conditions
- Aggregate revenue and volume without hiding bad data
- Validate customer and product relationships
On this page
Dataset shape
The file contains 250 data rows and seven columns.
| Column | Meaning |
|---|---|
order_id |
Order identifier; two complete records are intentionally duplicated |
customer_id |
Customer key, including two unmatched values |
order_date |
ISO date covering eight months of 2025 |
product_id |
Product key, including one intentional bad reference |
quantity |
Units ordered, with one blank value |
unit_price |
Transaction price, with one blank value and controlled discounts |
status |
Completed, Pending, Cancelled, or Returned |
What to practice
This is the main transactional dataset for the series. Use it for boolean filtering, sorting, nullable numeric conversion, calculated revenue, status counts, time grouping, customer aggregation, and joins. Repeated customers and products create useful groups, while cancelled and returned orders make business rules important: gross order value is not automatically recognized revenue.
Two rows use customer IDs that do not exist in customers.csv, one row references a nonexistent product, and two records are duplicated. Those conditions support join validation and reconciliation exercises. Missing quantity and price values force an explicit decision instead of silently producing misleading totals.
The data is deterministic, so examples and expected counts remain stable across downloads and builds. Treat each imperfection as a controlled test case and document how an analysis handles it.
Go deeper: related articles
Python
Filtering Rows in pandas
Build readable, null-safe pandas filters for realistic order questions and validate the result instead of relying on fragile expressions.
Python
GroupBy and Aggregations
Build trustworthy pandas summaries by defining grain, revenue rules, null behavior and joins before grouping real order data.
Python
Joining and Merging DataFrames
Join pandas DataFrames safely with explicit cardinality, unmatched-key checks, indicators and row-count reconciliation.
Python
Reading CSV Files with pandas
Load a real CSV into pandas deliberately, validate its shape and types, and avoid the ingestion assumptions that cause downstream data problems.