All resources

Dataset Python

Customer Dataset

A synthetic customer master dataset for practicing CSV ingestion, inspection, cleaning, filtering and joins with pandas.

Type
Dataset
Level
Beginner
Updated

In short

One hundred customer rows with countries, signup dates, segments, nullable emails, inconsistent casing and one intentional duplicate.

Who it is for

  • Data engineers learning practical pandas workflows
  • Readers practicing data-quality checks and joins

What it helps you do

  • Load and inspect a typed CSV dataset
  • Detect nulls, inconsistent categories and duplicates
  • Join customer attributes to transactional data

Dataset shape

The file contains 100 data rows and six columns. All names, organizations, and email addresses are synthetic; the reserved example.test domain cannot represent a real mailbox.

Column Meaning
customer_id Customer key used by the other learning datasets
customer_name Synthetic display name
country Country label, including three deliberate casing inconsistencies
signup_date ISO-formatted signup date
segment Enterprise, Small Business, Consumer, or Public Sector
email Synthetic email, blank in a few rows

What to practice

Use this dataset to practice read_csv, head, info, column selection, boolean filters, date parsing, normalization, missing-value checks, and duplicate detection. One entire customer row is duplicated deliberately, so duplicated() and drop_duplicates() have a known result. Three emails are blank, and three country values use inconsistent casing.

The identifiers are designed to join to orders.csv and website_events.csv. Most foreign keys match, but those datasets also contain unmatched or anonymous activity. That makes inner, left, and anti-join checks produce meaningfully different results rather than perfect classroom output.

Keep an untouched copy of the download and perform cleaning in a DataFrame. The imperfections are teaching fixtures, not accidental corruption.