12.1 Introduction to Sets
What a set is, its core characteristics, why automatic deduplication and O(1) average membership testing are the entire reason it exists as a distinct type, and where sets fit against lists and tuples.
What Is a Set?
What Is It?
An unordered collection of unique, hashable values. Written with curly braces, or via set().
>>> unique_ports = {80, 443, 8080}
>>> unique_ports
{80, 443, 8080}
Sets were already introduced briefly in 5.6 Set Data Types; this chapter covers them in full.
Why Does It Matter?
Whenever “does this exist?” or “remove duplicates” matters more than order or position, a set is the right tool — this is a genuinely different job from what a list or tuple is built for.
Characteristics of Sets
- Unordered — no guaranteed position or insertion order (see 12.6 Set Indexing and Ordering).
- No duplicates — adding an existing value is a no-op.
- Mutable — elements can be added/removed after creation (see 12.4 Adding Elements and 12.5 Removing Elements).
- Elements must be hashable — no lists or dicts as members (see 12.17 Common Mistakes); this follows the same rule covered in 5.9 Mutable vs Immutable Types.
Why Use Sets?
Two things a set does better than any other built-in type: automatic deduplication, and near-instant membership testing — O(1) average, versus a list’s O(n) (see 10.6 List Operators for the list-side comparison). Both are covered in depth in 12.13 Performance.
Real-World Applications
- Deduplicating a list of email addresses
- Tracking which users have already been notified
- Finding common tags between two articles
- Validating a value against an allowed set of options
Sets in DevOps
Deduplicating IP addresses from logs, comparing installed vs. required packages, diffing security group rules — 12.16 Sets in DevOps is dedicated entirely to these patterns.
>>> unique_ips = {"10.0.0.1", "10.0.0.2", "10.0.0.1"}
>>> unique_ips
{'10.0.0.1', '10.0.0.2'}
Quick Interview Answer
“A set is an unordered collection of unique, hashable values, built for exactly two jobs: automatic deduplication and fast membership testing. Internally it’s a hash table (see 12.3 Internal Representation), so checking
x in sis O(1) average instead of the O(n) a list requires — that performance difference is the entire reason the type exists. The trade-offs that follow directly from being hash-based: no indexing, no guaranteed iteration order, and every element must itself be hashable.”
Common Mistakes
- Reaching for a list and manually checking
inrepeatedly, when converting to a set once up front would make each check O(1) instead of O(n) (see 12.18 Best Practices). - Assuming a set preserves insertion order the way a list, tuple, or (since Python 3.7) a dict does — it doesn’t.
- Writing
s = {}expecting an empty set — that creates an emptydictinstead (see 12.2 Creating Sets).
Add More Questions to This Guide
Know a question that should be here? Share it and help the community!
Open Google Form