Python Fundamentals

Choose a chapter or lesson from the contents. Read one lesson at a time, then use Previous Lesson or Next Lesson to continue.

Introduction to Python

1.1 What is Python?

Python is a programming language. You can use it to write small scripts, automation tools, web applications, and many other types of software.

Python is popular because its code is easy to read and quick to write.

print("Hello, Python!")

When you run this program, Python displays:

Hello, Python!

A Python program is usually saved in a file ending with .py. You can run the file with the Python interpreter:

python3 hello.py

1.2 A Short History

Python was created by Guido van Rossum. The first public version was released in 1991.

The language was designed with a simple goal: developers should be able to write clear and useful programs without unnecessary complexity.

Python 3 is the current version. Python 2 reached the end of its life on January 1, 2020, so new projects should use Python 3.

# Python 3
print("Use Python 3 for new projects")

1.3 Main Features of Python

  1. Easy to read syntax
  2. Cross platform
  3. Large standard library
  4. Many third-party packages
  5. Dynamically typed

1. Easy-to-Read Syntax

Python code often looks close to plain English. It uses indentation to show which lines belong together.

name = "Alex"

if name:
    print(f"Hello, {name}")

2. Cross-Platform

Python is available on Linux, Windows, and macOS. The same script can often run on all three operating systems.

3. Large Standard Library

Python includes useful modules for files, JSON, dates, networking, and many other common tasks. You can use these modules without installing extra packages.

import json

server = {"name": "web-01", "status": "running"}
print(json.dumps(server))

4. Many Third-Party Packages

When the standard library is not enough, you can install packages from PyPI. Examples include:

  • boto3 for AWS
  • requests for HTTP APIs
  • pytest for testing
  • pandas for data analysis

5. Dynamically Typed

You do not have to declare a variable type before using it. Python works out the type from the value.

count = 5
name = "web-server"

print(type(count))
print(type(name))

1.4 Why Learn Python?

Python is useful because it helps you solve real problems with a small amount of code.

  • Automation: Automate repetitive tasks such as moving files, creating reports, and checking servers.
  • DevOps: Write scripts for CI/CD, monitoring, deployments, and infrastructure management.
  • Cloud: Use cloud SDKs such as boto3 to manage AWS resources.
  • Web development: Build websites and APIs with frameworks such as Django, Flask, and FastAPI.
  • Data and AI: Work with data and machine learning using libraries such as pandas, NumPy, and PyTorch.

1.5 Where Is Python Used?

flowchart TD P((Python)) P --- WEB[Web Development] P --- AUTO[Automation] P --- DS[Data Science] P --- ML[Machine Learning] P --- SCRIPT[Scripting] P --- NET[Networking] P --- SEC[Cybersecurity]

Python is commonly used for:

  1. Web development: Building websites and APIs.
  2. Automation: Replacing repetitive manual work with scripts.
  3. Data science: Cleaning, analyzing, and visualizing data.
  4. Machine learning: Building and using machine-learning models.
  5. Networking: Creating API clients, monitoring tools, and network scripts.
  6. Cybersecurity: Analyzing logs and automating security checks.

1.6 Why Python Is Useful for DevOps

DevOps engineers work with servers, cloud services, APIs, containers, and CI/CD pipelines. Python can connect these systems and automate work between them.

Infrastructure Automation

Python can connect to servers and run administrative tasks. Libraries such as paramiko can be used for SSH automation.

import paramiko

client = paramiko.SSHClient()
client.set_missing_host_key_policy(paramiko.AutoAddPolicy())
client.connect('server01', username='deploy', key_filename='id_rsa')
stdin, stdout, stderr = client.exec_command('uptime')
print(stdout.read().decode())

AWS Automation

boto3 is the AWS SDK for Python. It allows Python programs to work with AWS services such as EC2, S3, Lambda, and IAM.

import boto3

ec2 = boto3.client('ec2')
response = ec2.describe_instances()
# The response contains information about EC2 instances.

CI/CD

Python scripts can run tests, build packages, check deployments, and trigger other tools inside Jenkins, GitHub Actions, or GitLab CI.

Monitoring

Python can check application health, read metrics, process logs, and send alerts.

Configuration Management

Ansible is built with Python. Python knowledge is useful when creating custom Ansible modules or automation scripts.

1.7 Python Compared with Other Tools

flowchart LR subgraph Python direction TB P1[General-purpose] P2[Readable] P3[Huge ecosystem] end subgraph Bash direction TB B1[Best for quick] B2[shell/OS-level] B3[glue scripts] end subgraph Go direction TB G1[Compiled, fast] G2[Great for CLI] G3[tools & binaries] end subgraph PowerShell direction TB PS1[Deep Windows /] PS2[Azure integration] end

Python and Bash

Bash is excellent for short commands and simple Linux tasks. Python is usually better when a script needs data structures, reusable functions, API calls, or detailed error handling.

Python and Go

Go is compiled and is a good choice for fast, standalone tools. Python is often quicker to write and is popular for internal automation.

Python and PowerShell

PowerShell is a strong choice for Windows and Microsoft administration. Python is a good choice when the same automation must run across Linux, Windows, and macOS.

When to Choose Each?

ScenarioBest Choice
TaskA good choice
——
Short Linux command or pipelineBash
Automation with files, APIs, or dataPython
Fast standalone command-line toolGo
Windows administrationPowerShell
Cross-platform cloud automationPython

Quick Interview Answer

Python is a general-purpose programming language created by Guido van Rossum and first released in 1991. It is popular because its syntax is easy to read and it helps developers write useful programs quickly. Python is widely used for automation, DevOps, cloud engineering, web development, data science, and machine learning.

Key Takeaways

  • Python is a readable, general-purpose programming language.
  • Python 3 should be used for new projects.
  • Python works on Linux, Windows, and macOS.
  • Python is useful for automation, cloud, DevOps, web development, and data work.
  • Python can be extended with packages from PyPI.

Common Mistakes

  • Assuming Python is only used for data science or machine learning.
  • Using Python 2 for a new project. Python 2 is no longer supported.
  • Choosing Python for every task. Bash, Go, or PowerShell may be a better fit in some situations.

Chapter practice checkpoint

Run python --version (or python3 --version on systems using that command). Save print("Hello, automation!") in hello.py, then run it from the terminal with python hello.py.

Expected output: Hello, automation! on its own line.

Knowledge check: What is the difference between the interpreter prompt and a saved script?

Check your answerThe prompt evaluates input interactively. A script stores statements in a file that the interpreter executes, making the work repeatable and shareable.

You are ready to continue when you can locate your script, run it, change its message, and explain the output.

Installation of Python

1.1 What Do You Need?

To start writing Python, you need:

  1. Python 3 installed on your computer.
  2. A terminal to run commands.
  3. A text editor such as VS Code.

You can download Python from python.org.

1.2 Install Python on Windows

  1. Open python.org/downloads.
  2. Download the latest Python 3 installer for Windows.
  3. Start the installer.
  4. Select Add python.exe to PATH.
  5. Select Install Now.

The PATH option allows you to run Python by typing python in a terminal.

Open PowerShell and check the installation:

python --version

You can also use the Windows Python launcher:

py --version

1.3 Install Python on Ubuntu or Debian

Open a terminal and run:

sudo apt update
sudo apt install python3 python3-pip python3-venv

Check the installation:

python3 --version
pip3 --version

1.4 Install Python on Fedora, RHEL, or CentOS

Run:

sudo dnf install python3 python3-pip

Then check the installation:

python3 --version

1.5 Install Python on macOS

Homebrew is a common way to install Python on macOS. If Homebrew is already installed, run:

brew install python

Check the installation:

python3 --version
pip3 --version

1.6 Verify That Python Works

Run a small Python command:

python3 -c "print('Python is working')"

On Windows, use python or py if python3 is not available:

python -c "print('Python is working')"

You should see:

Python is working

1.7 Write Your First Python File

Create a file named hello.py:

print("Hello, Python!")

Run it from the terminal:

python3 hello.py

On Windows, use:

python hello.py

The output should be:

Hello, Python!

1.8 Choose an Editor

You can write Python in any text editor, but these tools make learning easier:

  • VS Code: A lightweight editor with Python support, debugging, and extensions.
  • PyCharm: A full Python development environment.
  • IDLE: A simple editor included with many Python installations.
  • Jupyter Notebook: Useful for learning, experiments, and data analysis.

For this course, VS Code with the official Python extension is a good choice.

1.9 Create a Virtual Environment

A virtual environment keeps the packages for one project separate from other projects.

Create one inside your project folder:

python3 -m venv .venv

Activate it on Linux or macOS:

source .venv/bin/activate

Activate it on Windows PowerShell:

.venv\Scripts\Activate.ps1

After activation, your terminal usually shows (.venv). You can now install packages without affecting other projects.

To leave the virtual environment, run:

deactivate

1.10 Install a Python Package

pip is Python’s package installer. For example, install the requests package inside an activated virtual environment:

python -m pip install requests

Check that it works:

python -c "import requests; print(requests.__version__)"

Using python -m pip makes sure that pip belongs to the Python interpreter you are using.

1.11 What Happens When You Run Python?

When you run a Python file, Python reads the file and executes the instructions.

flowchart LR A[Python file] --> B[Python interpreter] B --> C[Program output]

For now, you only need to remember this:

  • You write Python code in a .py file.
  • The Python interpreter runs the file.
  • Python displays the result or an error message.

Python may create a __pycache__ folder while it runs. This folder contains temporary cached files. You normally do not need to edit it.

1.12 Basic Python Style

Python uses indentation to group code. Use four spaces for each indentation level.

name = "Alex"

if name:
  print(f"Hello, {name}")

Use clear names:

server_count = 3

Avoid unclear names:

x = 3

Python’s common style guide is called PEP 8. It helps teams write code that looks consistent and is easy to read.

1.13 Common Installation Problems

Python Command Not Found

This usually means Python is not installed or its directory is not in the PATH.

On Windows, reinstall Python and select Add python.exe to PATH. Then close and reopen PowerShell.

On Linux or macOS, try python3 instead of python.

The Wrong Python Version Appears

Several Python versions can exist on one computer. Check the version before installing packages or running a project:

python3 --version

On Windows:

py --list

pip Installs into the Wrong Python

Use this form instead of calling pip by itself:

python -m pip install package-name

PowerShell Does Not Allow Virtual Environment Activation

If PowerShell blocks the activation script, use Command Prompt, or ask an administrator to review the PowerShell execution policy for your machine. Do not change security settings without understanding your organization’s policy.

1.14 Practice Exercise

  1. Install Python 3.
  2. Confirm the Python version.
  3. Create a file named system_check.py.
  4. Add this code:
import platform

print(f"Operating system: {platform.system()}")
print(f"Python version: {platform.python_version()}")
  1. Run the file from your terminal.

Interview Questions

  • How do you check whether Python is installed?
  • What is the difference between python, python3, and py?
  • What is PATH, and why is it important?
  • What is a virtual environment?
  • Why should you use python -m pip?
  • How do you run a Python file from the terminal?

Quick Interview Answer

To install Python, I install Python 3 from the official Python website or the operating system package manager. I verify it with python --version or python3 --version, create a virtual environment for each project, and install packages with python -m pip. On Windows, I make sure that Python is added to PATH during installation.

Key Takeaways

  • Install Python 3, not Python 2.
  • Always verify the installation from a terminal.
  • Use a virtual environment for each project.
  • Use python -m pip to install packages for the correct interpreter.
  • Use VS Code or another editor to write Python files.

Chapter practice checkpoint

Run python --version (or python3 --version on systems using that command). Save print("Hello, automation!") in hello.py, then run it from the terminal with python hello.py.

Expected output: Hello, automation! on its own line.

Knowledge check: What is the difference between the interpreter prompt and a saved script?

Check your answerThe prompt evaluates input interactively. A script stores statements in a file that the interpreter executes, making the work repeatable and shareable.

You are ready to continue when you can locate your script, run it, change its message, and explain the output.

Write and Run Your First Python Program

Write and Run Python Program

1.1 Create Your First Python File

Open VS Code, IDLE, PyCharm, or another text editor. Create a file named hello.py and add this code:

print("Hello, World!")

Save the file. The .py ending tells your editor that the file contains Python code.

1.2 Understand the Program

The line contains two parts:

  • print() is a Python function that displays text.
  • "Hello, World!" is the text we want to display.

The output is:

Hello, World!

Python does not require a semicolon at the end of this line.

1.3 Run the Program on Windows

  1. Open PowerShell or Command Prompt.
  2. Move to the folder where you saved hello.py.
cd C:\Users\yourname\projects
  1. Run the program:
python hello.py

You can also use the Windows Python launcher:

py hello.py

The output should be:

Hello, World!

1.4 Run the Program on macOS or Linux

  1. Open a terminal.
  2. Move to the folder where you saved hello.py.
cd ~/projects
  1. Run the program:
python3 hello.py

On macOS and Linux, python3 is usually the correct command for Python 3.

1.5 Run the Program in VS Code

  1. Open the hello.py file in VS Code.
  2. Select the Python interpreter if VS Code asks you to choose one.
  3. Click the Run Python File button in the top-right corner.

VS Code runs the file and shows the output in the terminal panel.

You can also right-click inside the file and select Run Python File in Terminal.

1.6 Run a Python Command Without Creating a File

You can use the -c option to run a short command directly.

On Windows:

python -c "print(2 + 2)"

On macOS or Linux:

python3 -c "print(2 + 2)"

The output is:

4

This is useful for quick checks, but larger programs should be saved in .py files.

1.7 Use the Python Interactive Shell

Run python on Windows or python3 on macOS/Linux without a file name:

>>> 2 + 3
5
>>> print("Testing Python")
Testing Python
>>> exit()

The interactive shell runs one line at a time. It is useful for testing small ideas.

1.8 Run a Script Directly on Linux or macOS

You can run a script directly by adding a shebang as the first line:

#!/usr/bin/env python3

print("Hello, World!")

Make the file executable:

chmod +x hello.py

Run it:

./hello.py

On Windows, run the file with python hello.py or py hello.py instead.

1.9 Add Variables and User Input

Now create a slightly more useful program:

name = input("What is your name? ")
print(f"Hello, {name}!")

When you run the program, Python waits for you to type a name and press Enter.

Example:

What is your name? Alex
Hello, Alex!

1.10 Common Mistakes

Running from the Wrong Folder

If Python cannot find the file, move to the correct folder with cd or provide the full file path.

can't open file 'hello.py': No such file or directory

Using Different Quote Characters

The opening and closing quotes must match:

print("Hello")

This is incorrect:

print("Hello')

Saving the Wrong File Name

Make sure the file is saved as hello.py, not hello.py.txt. On Windows, enable File name extensions in File Explorer if you cannot see the full file name.

Using the Wrong Python Command

Use the command that matches your operating system:

Operating systemCommand
Windowspython hello.py or py hello.py
macOSpython3 hello.py
Linuxpython3 hello.py

Forgetting to Save the File

Save the file before running it. Otherwise, Python may run an older version of your code.

1.11 Practice Exercise

Create a file named system_check.py:

import platform

name = input("Enter your name: ")
print(f"Hello, {name}!")
print(f"Operating system: {platform.system()}")
print(f"Python version: {platform.python_version()}")

Run the program and record the output.

Interview Questions

  • How do you create a Python program?
  • How do you run a Python file on Windows?
  • How do you run a Python file on Linux or macOS?
  • What is the difference between python, python3, and py?
  • What does the print() function do?
  • What is a shebang line?
  • What is the Python interactive shell?

Quick Interview Answer

I create a file with a .py extension, write Python code in it, and run it from a terminal. On Windows I usually use python file.py or py file.py. On Linux and macOS I usually use python3 file.py. I can also run the file from an editor such as VS Code.

Key Takeaways

  • Python programs are usually saved in .py files.
  • print() displays text or values.
  • Use python or py on Windows and python3 on macOS/Linux.
  • The terminal must be in the correct folder before you run a file.
  • VS Code can run Python files directly.

Chapter practice checkpoint

Run python --version (or python3 --version on systems using that command). Save print("Hello, automation!") in hello.py, then run it from the terminal with python hello.py.

Expected output: Hello, automation! on its own line.

Knowledge check: What is the difference between the interpreter prompt and a saved script?

Check your answerThe prompt evaluates input interactively. A script stores statements in a file that the interpreter executes, making the work repeatable and shareable.

You are ready to continue when you can locate your script, run it, change its message, and explain the output.

4.1 Introduction to Python Syntax

Python Syntax

1.1 What Is Python Syntax?

Syntax means the rules for writing valid Python code. Just as a sentence needs correct grammar, a Python program needs correct syntax.

For example, this is valid Python:

print("Python is easy to read")

Python understands the print() function and the text inside the quotes.

1.2 Why Does Syntax Matter?

Python checks your syntax before it runs the program. If the syntax is incorrect, Python stops and shows an error.

print("Hello")

This code is valid. The following code is missing a closing parenthesis:

print("Hello"

Python reports a SyntaxError because it cannot understand the line.

1.3 Python Reads Code from Top to Bottom

Python normally runs statements in order, starting with the first line:

print("Step 1")
print("Step 2")
print("Step 3")

The output is:

Step 1
Step 2
Step 3

This makes the order of your code important.

flowchart TD A[Read first line] --> B[Read next line] B --> C[Run the instruction] C --> D{More lines?} D -->|Yes| B D -->|No| E[Program ends]

1.4 Python Uses Indentation

Indentation means spaces at the beginning of a line. Python uses indentation to show which lines belong inside a block.

age = 20

if age >= 18:
  print("You are an adult")

The indented print() line belongs to the if statement.

Use four spaces for each indentation level:

if True:
  print("This line is inside the if block")

Do not use random numbers of spaces:

if True:
  print("This may cause an indentation problem")

Most editors can insert four spaces automatically when you press Tab.

1.5 Colons Start Code Blocks

Statements such as if, for, while, and function definitions use a colon at the end of the first line:

if 10 > 5:
  print("The condition is true")

The colon tells Python that an indented block is coming next.

1.6 Comments

Comments are notes for people reading the code. Python ignores comments when it runs the program.

# Check whether the server is healthy
server_is_healthy = True
print(server_is_healthy)

Use comments to explain why something is done when the code is not obvious.

1.7 Names and Values

You can store a value in a variable by giving it a name:

server_name = "web-01"
server_port = 8080

print(server_name)
print(server_port)

Use clear names with lowercase letters and underscores:

retry_count = 3

Avoid unclear names:

x = 3

Python names cannot contain spaces and cannot start with a number.

1.8 Syntax Errors and Logic Errors

Syntax Error

A syntax error means Python cannot understand the code. The program stops before it starts:

if True
  print("Hello")

This code is missing a colon.

Logic Error

A logic error means the code is valid, but the result is wrong:

price = 100
discount = 20

final_price = price + discount  # The intended operation was probably subtraction.
print(final_price)

Python runs this code because the syntax is valid. It cannot know what result you intended.

1.9 Syntax Error and Runtime Error

A runtime error happens while the program is running:

server_count = 3
print(server_count + " servers")

Python cannot add a number and text, so it raises an error when it reaches that line.

Remember:

Error typeMeaning
SyntaxErrorPython cannot understand the code structure
Runtime errorA problem happens while the code is running
Logic errorThe code runs but gives the wrong result

1.10 Practice Exercise

Create a file named syntax-practice.py:

name = "Alex"
age = 25

if age >= 18:
  print(f"{name} is an adult")
else:
  print(f"{name} is a minor")

Run the file and then change the value of age. Observe how the output changes.

Interview Questions

  • What is syntax in Python?
  • Why is indentation important in Python?
  • What does a colon mean after an if statement?
  • What is the difference between a syntax error and a logic error?
  • What is the difference between a runtime error and a syntax error?
  • How should Python variables be named?

Quick Interview Answer

Python syntax is the set of rules used to write valid Python code. Python uses indentation to define code blocks, and statements such as if and for use a colon before an indented block. A syntax error stops the program before it runs, while a logic error allows the program to run but produces the wrong result.

Key Takeaways

  • Syntax is the grammar of Python code.
  • Python reads code from top to bottom.
  • Indentation defines blocks of code.
  • Colons start blocks after statements such as if and for.
  • Comments begin with #.
  • Syntax errors stop a program before it runs.
  • Logic errors produce incorrect results without necessarily showing an error.

4.2 Structure of a Python Program

1.1 What Is the Structure of a Python Program?

The structure of a Python program means the way the code is organized in a file.

A small Python file may contain only one line:

print("Hello")

As the program becomes larger, organizing the code makes it easier to read, test, and update.

1.2 A Common Python File Layout

Many Python files follow this order:

  1. Imports
  2. Constants and configuration
  3. Functions
  4. Classes, when needed
  5. The main program
flowchart TD A["Imports\nimport os\nimport sys"] --> B["Constants\nMAX_RETRIES = 3"] B --> C["Function / class definitions\ndef main():\n ..."] C --> D["Entry-point guard\nif __name__ == '__main__':\n main()"]

1.3 Imports

An import loads code that you want to use. In this example, platform is a standard Python module:

import platform

print(platform.system())

Imports are usually placed at the top of the file so readers can quickly see what the program needs.

1.4 Constants and Configuration

Constants are values that normally do not change while the program runs. Python developers commonly write constant names in uppercase:

MAX_RETRIES = 3
DEFAULT_REGION = "us-east-1"

Do not store passwords, API keys, or other secrets directly in a Python file.

1.5 Functions

A function is a named group of instructions. It runs when you call it.

def greet_user(name):
    print(f"Hello, {name}!")


greet_user("Alex")

Defining a function does not run its instructions immediately. The function runs only when you call it.

1.6 The main() Function

The main() function usually contains the main steps of a script:

def main():
    print("Checking the server")


main()

Using a main() function keeps the starting point of the program clear.

1.7 The if __name__ == "__main__": Check

This common Python line means: run the following code only when this file is started directly.

def main():
    print("The script is running")


if __name__ == "__main__":
    main()

The check is useful because another Python file can import your functions without automatically starting the whole script.

1.8 How Python Runs the File

Python reads top-level code from top to bottom. It creates functions when it reaches their definitions, then runs the code that calls them.

print("1. Start")


def show_message():
    print("2. Inside the function")


print("3. Before the function call")
show_message()

The output is:

1. Start
3. Before the function call
2. Inside the function

1.9 Write Readable Python

Good structure helps other people understand your code.

  • Use clear names such as server_count instead of x.
  • Keep functions small and focused.
  • Use four spaces for indentation.
  • Put imports at the top.
  • Keep configuration separate from the main logic.

1.10 Practice Exercise

Create a file named server_check.py:

import platform


def show_server_info():
    print(f"Operating system: {platform.system()}")
    print("Server check complete")


def main():
    show_server_info()


if __name__ == "__main__":
    main()

Run it with:

python3 server_check.py

On Windows, use:

python server_check.py

Quick Interview Answer

A Python file commonly contains imports, configuration values, functions, classes when needed, and a main program section. The if __name__ == "__main__": check runs the main function only when the file is executed directly. This allows the same file to work as both a script and an importable module.

Key Takeaways

  • Organize Python code into clear sections.
  • Put imports near the top of the file.
  • Use functions to group reusable work.
  • Use main() as the clear starting point of a script.
  • Use the __main__ check to prevent code from running during import.
  • Clear structure makes Python scripts easier to read and test.

4.3 Statements

1.1 What Is a Statement?

A statement is an instruction that tells Python to do something.

name = "Alex"
print(name)

The first statement stores a name. The second statement displays it.

Python normally runs statements from top to bottom.

1.2 Simple Statements

A simple statement usually fits on one line and performs one action.

server_count = 3              # assignment
print(server_count)            # function call
import platform                # import

Other common simple statements include return, break, and continue inside functions or loops.

1.3 Compound Statements

A compound statement controls a block of code. Its first line ends with a colon, and the lines below it are indented.

server_status = "running"

if server_status == "running":
    print("The server is healthy")

The if line is the header. The indented print() line is the body.

Common compound statements include:

  • if for decisions
  • for and while for loops
  • def for functions
  • class for classes
  • try for error handling

1.4 Statements with Conditions

An if statement runs its block only when the condition is true:

disk_usage = 70

if disk_usage > 80:
    print("Disk usage is high")
else:
    print("Disk usage is normal")

The else block runs when the if condition is false.

1.5 Statements with Loops

A for statement repeats a block for each item:

servers = ["web-01", "web-02", "web-03"]

for server in servers:
    print(server)

The output is:

web-01
web-02
web-03

1.6 Multiple Statements on One Line

Python allows multiple simple statements on one line with semicolons:

first_name = "Alex"; print(first_name)

However, this style is harder to read. Write one statement per line instead:

first_name = "Alex"
print(first_name)

One statement per line is the usual Python style and makes errors easier to find.

1.7 Long Statements on Multiple Lines

Long statements can be split across lines inside brackets:

server_names = [
    "web-01",
    "web-02",
    "web-03",
]

Python understands that the list continues until the closing bracket.

You can also use parentheses for a long calculation:

total_cost = (
    monthly_compute_cost
    + monthly_storage_cost
    + monthly_network_cost
)

This is easier to read than one very long line.

Avoid using a backslash for line continuation when brackets can be used:

total = (
    first_value
    + second_value
)

1.8 The pass Statement

The pass statement does nothing. It is useful when you want to create a block and add the real code later:

def check_server():
    pass

Without pass, Python would report an error because the function has no body.

1.9 A Complete Example

This example combines assignments, a condition, a function, and a loop:

servers = ["web-01", "web-02"]
healthy_servers = 0


def show_server(server):
    print(f"Checking {server}")


for server in servers:
    show_server(server)
    healthy_servers += 1


print(f"Healthy servers: {healthy_servers}")

1.10 Practice Exercise

Create a program that checks a list of server statuses:

server_statuses = ["running", "stopped", "running"]

for status in server_statuses:
    if status == "running":
        print("Server is healthy")
    else:
        print("Server needs attention")

Change the list and observe how the output changes.

Interview Questions

  • What is a statement in Python?
  • What is the difference between a simple and compound statement?
  • Why does a compound statement need a colon?
  • Why is indentation important after a compound statement?
  • Can you place multiple statements on one line?
  • How can you split a long statement across multiple lines?

Quick Interview Answer

A statement is an instruction for Python to execute. A simple statement usually performs one action on one line, such as an assignment or function call. A compound statement, such as if, for, or def, ends with a colon and contains an indented block. Python allows semicolons, but one statement per line is easier to read.

Key Takeaways

  • A statement is an instruction for Python.
  • Simple statements usually fit on one line.
  • Compound statements contain an indented block.
  • Compound statement headers end with a colon.
  • Prefer one statement per line.
  • Use brackets or parentheses to split long statements. 1 2

## Line Continuation

### What Is It?

Splitting one logical statement across multiple physical lines, either explicitly with a trailing backslash or implicitly inside brackets.

### Why Is It Used?

It keeps long expressions readable instead of one very long line.

### How Is It Used?

See [4.15 Multiple Statements & Line Continuation](../15-multiple-statements-and-line-continuation/) for the full comparison of both styles.

```python
total = 1 + \
        2 + \
        3
>>> total
6

Quick Interview Answer

“A simple statement fits on one line and does one thing — an assignment, a function call, an import. A compound statement has a colon-terminated header and an indented block beneath it — if, for, def, class. You can technically cram multiple simple statements onto one line with semicolons, but PEP 8 discourages it.”

Common Mistakes

  • Forgetting the colon at the end of a compound statement’s header line — the single most common beginner SyntaxError.
  • Overusing semicolons to combine simple statements, which hurts readability without any real performance benefit.

4.4 Indentation

1.1 What Is Indentation?

Indentation means spaces at the beginning of a line. Python uses indentation to show which lines belong together.

In many languages, braces mark a block. Python uses indentation instead:

if temperature > 30:
    print("It is hot")

The indented print() line belongs to the if statement.

1.2 The Basic Rule

When a line ends with a colon, the next related lines must be indented:

if server_status == "running":
    print("Server is running")

Statements at the same level should have the same indentation:

if server_status == "running":
    print("Server is running")
    print("Health check passed")

1.3 Use Four Spaces

The usual Python style is four spaces for each indentation level:

if True:
    print("Level one")

Most code editors insert four spaces when you press Tab. Configure your editor to use spaces instead of tabs for Python files.

1.4 Indentation in Functions and Loops

Indentation is used with functions, loops, conditions, and other blocks:

def check_servers(servers):
    for server in servers:
        print(f"Checking {server}")

The indentation shows the relationship:

def check_servers        level 0
    for server           level 1
        print(...)       level 2

1.5 Nested Blocks

A nested block is a block inside another block. Add one indentation level for each new block:

servers = ["web-01", "web-02"]

for server in servers:
    if server == "web-01":
        print(f"Checking important server: {server}")
    else:
        print(f"Checking server: {server}")

The if statement is inside the for loop, so it has more indentation.

flowchart TD A[for loop] --> B[if condition] B --> C[Indented action]

1.6 Incorrect Indentation

This code causes an error because the line after the colon is not indented:

if True:
print("This is not inside the if block")

Python reports an IndentationError.

Correct it by adding four spaces:

if True:
    print("This is inside the if block")

1.7 Inconsistent Indentation

Lines in the same block must use the same indentation:

if True:
    print("First line")
  print("Second line")

The second line does not line up with the first line, so Python may report an IndentationError.

1.8 Do Not Mix Tabs and Spaces

Mixing tabs and spaces can cause a TabError or make the code difficult to read.

Use your editor’s Python formatting settings to insert spaces. In VS Code, you can use the indentation indicator in the bottom status bar to convert indentation to spaces.

1.9 Indentation with else

The else statement must line up with its matching if:

disk_usage = 70

if disk_usage > 80:
    print("Disk usage is high")
else:
    print("Disk usage is normal")

The code inside each branch is indented, but if and else are at the same level.

1.10 Practice Exercise

Create a file named indentation-practice.py:

services = ["api", "database", "cache"]

for service in services:
    if service == "database":
        print(f"Checking critical service: {service}")
    else:
        print(f"Checking service: {service}")

Change the indentation intentionally and observe the error. Then fix the indentation and run the program again.

Interview Questions

  • Why is indentation important in Python?
  • How many spaces are normally used for one indentation level?
  • What happens when indentation is missing after a colon?
  • What is a nested block?
  • Why should tabs and spaces not be mixed?
  • How should else align with if?

Quick Interview Answer

Python uses indentation to define blocks of code instead of braces. Lines in the same block must have the same indentation, and a new block normally uses four spaces. Indentation is required after statements such as if, for, and def, and mixing tabs and spaces can cause an error.

Key Takeaways

  • Indentation is part of Python syntax.
  • Use four spaces for each indentation level.
  • Indent code after a line ending with :.
  • Keep statements in the same block aligned.
  • Use more indentation for nested blocks.
  • Do not mix tabs and spaces.

4.5 Comments

1.1 What Is a Comment?

A comment is a note written for people reading the code. Python ignores comments when it runs the program.

Comments are useful when they explain a decision, a warning, or a part of the code that is not obvious.

flowchart LR A[Python file] --> B{Line starts with #?} B -->|Yes| C[Python ignores the comment] B -->|No| D[Python reads and executes the code]

1.2 Single-Line Comments

Start a comment with #. Everything after # on that line is ignored:

# Check the API three times because the service may respond slowly.
max_retries = 3

You can also place a short comment after code:

timeout_seconds = 30  # Keep the request from waiting forever.

Use an end-of-line comment only when it remains easy to read.

1.3 Comments Before Code

A comment can explain the purpose of the next few lines:

# Skip temporary files before uploading the directory to S3.
for file_name in files:
    if not file_name.endswith(".tmp"):
        upload(file_name)

The comment explains why the condition exists. The code already shows what the condition does.

1.4 Multi-Line Comments

Python does not have a special multi-line comment symbol. For a longer comment, use # on each line:

# This script checks the health of each service.
# It prints a warning when a service is not running.
# The output can be used by a monitoring job.

This is the clearest and most reliable way to write a multi-line comment.

1.5 Comments and Triple Quotes Are Different

Triple quotes create a string, not a real multi-line comment:

"""
This looks like a comment,
but Python treats it as a string.
"""

An unused string may appear to work like a comment, but it is not the correct general-purpose comment style. Use # for comments.

1.6 What Is a Docstring?

A docstring is text that documents a module, function, class, or method. It is written inside triple quotes as the first statement in that object.

def check_server(server_name):
    """Return the health status of a server."""
    return f"{server_name} is healthy"

Unlike a normal comment, a docstring is available while the program runs:

print(check_server.__doc__)

Output:

Return the health status of a server.

Python’s help() function and many editors can display docstrings:

help(check_server)

1.7 Function Docstrings

Use a function docstring to explain what the function does, its inputs, and its result when that information is useful:

def calculate_retry_delay(attempt, base_delay=2):
    """Return the delay before the next retry attempt."""
    return base_delay ** attempt

The function name and code show how it works. The docstring gives a short explanation for someone who wants to use it.

1.8 Module and Class Docstrings

A module docstring goes at the top of a Python file:

"""Tools for checking application health."""

import requests

A class docstring goes immediately inside the class:

class HealthChecker:
    """Check whether an application endpoint is responding."""

    pass

1.9 Good and Bad Comments

This comment does not add much information:

# Add one to count.
count = count + 1

The code is already clear. This comment is more useful:

# Count only successful responses so failed checks are reported separately.
successful_checks += 1

Good comments usually explain why, not only what.

1.10 Keep Comments Correct

Update a comment when the code changes. A comment that disagrees with the code can mislead someone:

# Retry five times.
for attempt in range(3):
    retry_request()

The comment and code disagree. Change one of them so they describe the same behavior.

1.11 Comments in DevOps Scripts

Comments are especially useful in automation scripts for explaining:

  • Why a retry or timeout value was chosen
  • Why a resource is skipped
  • What permission a cloud API call requires
  • Why a command is safe or potentially destructive
  • What format an input file must use
# The cleanup is limited to objects older than 30 days.
# Do not remove newer objects because they may be needed for recovery.
if object_age_days > 30:
    delete_object(object_name)

Never put passwords, access keys, tokens, or other secrets in comments. Comments are stored in the source code and may be shared with the repository.

1.12 Practice Exercise

Add useful comments and a docstring to this script:

def check_services(services):
    """Print the status of each service in the list."""
    for service in services:
        # Treat only running services as healthy.
        if service["status"] == "running":
            print(f"{service['name']} is healthy")
        else:
            print(f"{service['name']} needs attention")


services = [
    {"name": "api", "status": "running"},
    {"name": "worker", "status": "stopped"},
]

check_services(services)

Add one comment explaining a decision, then add a docstring to the function.

Interview Questions

  • What is a comment in Python?
  • How do you write a single-line comment?
  • How do you write a multi-line comment?
  • Are triple-quoted strings the same as comments?
  • What is a docstring?
  • Where should a function docstring be placed?
  • What is the difference between a comment and a docstring?
  • What makes a comment useful?

Quick Interview Answer

A Python comment starts with #, and Python ignores it when running the program. Python does not have a special multi-line comment syntax, so we use # on each line. A docstring is different: it is a triple-quoted string placed first inside a module, function, or class, and it can be viewed through __doc__ or help().

Key Takeaways

  • Use # to write comments.
  • Use comments to explain why code exists.
  • Use one # on each line for multi-line comments.
  • Use docstrings to document reusable functions, classes, and modules.
  • Keep comments short, clear, and up to date.
  • Never store secrets in source code comments.

4.6 Keywords

What Are Keywords?

What Is It?

Reserved words that are part of Python’s own syntax (if, for, def, class, …). The language grammar gives them special meaning, so they cannot be used as ordinary names.

Why Does It Matter?

Attempting to use a keyword as a variable name is a SyntaxError, not a warning.

>>> class = 5
  File "<stdin>", line 1
    class = 5
    ^^^^^
SyntaxError: invalid syntax

Complete Keyword List

Python 3.12 has 35 keywords, retrievable programmatically via the keyword module:

>>> import keyword
>>> keyword.kwlist
['False', 'None', 'True', 'and', 'as', 'assert', 'async', 'await',
'break', 'class', 'continue', 'def', 'del', 'elif', 'else', 'except',
'finally', 'for', 'from', 'global', 'if', 'import', 'in', 'is',
'lambda', 'nonlocal', 'not', 'or', 'pass', 'raise', 'return', 'try',
'while', 'with', 'yield']
>>> len(keyword.kwlist)
35
flowchart TD K["35 Python Keywords"] K --> CF["Control Flow\nif / elif / else\nfor / while\nbreak / continue / pass"] K --> FC["Functions & Classes\ndef / return / class\nlambda / yield"] K --> LV["Logical / Values\nand / or / not\nTrue / False / None\nin / is"] K --> EH["Error Handling & Scope\ntry / except / finally / raise\nglobal / nonlocal\nimport / from / as"]

Python’s 35 keywords, grouped by purpose.

Hard Keywords vs Soft Keywords

Hard Keywords

Hard keywords are always reserved. The parser recognizes them wherever an identifier could otherwise appear, so names such as class, for, and return are invalid in every context.

class = "deployment"  # SyntaxError: invalid syntax

Soft Keywords

Soft keywords are reserved only in specific grammar contexts. In Python 3.10 and later, match and case are soft keywords used by structural pattern matching. The underscore (_) also has special pattern-matching behavior. Python 3.12 adds type as a soft keyword for type alias statements.

match status:
  case 200:
    message = "ok"
  case _:
    message = "unexpected status"

# Outside a match statement, these names can still be used normally.
match = "rolling"
type = "service"

Use keyword.softkwlist when code must account for contextual keywords as well as the hard keyword list:

>>> import keyword
>>> keyword.softkwlist
['_', 'case', 'match', 'type']

The exact list is version-dependent. Check the interpreter that will run the code rather than hard-coding a count.

Keywords by Role

RoleKeywordsTypical purpose
Decisions and loopsif, elif, else, for, while, break, continue, passSelect paths and repeat work
Functions and generatorsdef, return, lambda, yieldDefine callable code and produce values lazily
Classes and objectsclass, del, is, inDefine types, remove bindings, and test identity or membership
Exceptionstry, except, else, finally, raise, assertDetect, handle, raise, and validate failures
Imports and aliasesimport, from, asLoad modules and assign local aliases
Context managerswithAcquire and release resources predictably
Scopeglobal, nonlocalRebind names outside the current local scope
Boolean and special valuesand, or, not, True, False, NoneCombine conditions and represent values
Asynchronous codeasync, awaitDefine and suspend asynchronous operations

Checking Names Safely

An identifier must pass both checks: it must have valid identifier syntax and must not be reserved by the interpreter.

import keyword


def is_available_name(candidate: str) -> bool:
  return candidate.isidentifier() and not keyword.iskeyword(candidate)


for candidate in ("service_name", "2services", "class"):
  print(candidate, is_available_name(candidate))

Expected output:

service_name True
2services False
class False

str.isidentifier() checks the shape of a name; keyword.iskeyword() checks whether the interpreter reserves it. Neither check prevents shadowing a built-in such as list or type, so use clear project-specific names and follow 4.16 Coding Standards (PEP 8).

Practical DevOps Example

Keywords frequently appear in automation scripts that inspect command results and handle failures:

import subprocess


def check_service(service_name: str) -> str:
  try:
    result = subprocess.run(
      ["systemctl", "is-active", service_name],
      check=True,
      capture_output=True,
      text=True,
    )
  except subprocess.CalledProcessError as error:
    raise RuntimeError(f"Unable to inspect {service_name}") from error
  else:
    return result.stdout.strip()
  finally:
    print(f"Finished checking {service_name}")

Here, def, try, except, raise, from, else, and finally are keywords that define the control flow. The service name remains data supplied to the function, not a dynamically constructed piece of Python syntax.

How Python Handles a Keyword Error

Python parses source code before executing it. If a hard keyword appears where an identifier is required, parsing stops and no statements from that module run.

flowchart LR S[Source code] --> P{Parser} P -->|valid grammar| C[Compile bytecode] C --> R[Run program] P -->|keyword used as a name| E[SyntaxError] E --> X[Execution stops]

This is why a keyword naming error cannot be fixed with try/except: the try block itself cannot execute until the source has successfully parsed.

Using help('keywords')

How Is It Used?

The interactive help system lists and explains every keyword directly from the interpreter, without needing internet access:

>>> help("keywords")
Here is a list of the Python keywords.  Enter any keyword to get more help.

False               class               from                or
None                continue            global              pass
...

Restrictions

Keywords cannot be used as variable, function, or class names. Use keyword.iskeyword() to check programmatically before using a name you’re unsure about:

>>> keyword.iskeyword("for")
True
>>> keyword.iskeyword("variable")
False

Keywords Are Not the Same as Built-ins

class is a keyword and cannot be assigned. list, type, and id are built-in names, not hard keywords, so Python permits them but shadowing them can break later code:

list = ["api", "worker"]
print(list("abc"))  # TypeError: 'list' object is not callable

Use a descriptive name such as services instead. This distinction is a common interview follow-up.

Troubleshooting Checklist

When Python reports a keyword-related SyntaxError:

  1. Read the caret (^) location, then inspect the complete statement around it.
  2. Check whether a variable, function, parameter, or class has the same spelling as a hard keyword.
  3. Run python -m py_compile your_script.py to catch parse errors without executing the script.
  4. Check the interpreter version with python --version; keyword behavior can differ between Python versions.
  5. Rename the identifier rather than trying to quote or escape it. Python has no escape syntax that turns a keyword into an ordinary identifier.

Quick Interview Answer

“Keywords are the 35 reserved words baked into Python’s grammar — if, for, def, class, import, and so on. They can never be used as an identifier; trying raises a SyntaxError immediately, not a runtime warning. You can list them at any time with keyword.kwlist or check a specific name with keyword.iskeyword().”

Common Mistakes

  • Trying to name a variable class, type, or list — class is a hard keyword (fails outright); type and list are built-in names, not keywords, so they’re technically legal but silently shadow the built-in.
  • Assuming the keyword count never changes across versions — soft keywords such as match, case, _, and type behave contextually and are reported separately by keyword.softkwlist.
  • Trying to catch a keyword SyntaxError at runtime — parsing happens before execution, so the source must be corrected first.
  • Treating a keyword as the same thing as a built-in — keywords are grammar tokens; built-ins are ordinary names initially provided by the built-in namespace.

Interview Follow-up Questions

  1. What is the difference between a keyword, a soft keyword, and a built-in name?
  2. How can you list Python keywords programmatically?
  3. Why can a try/except block not catch class = 1?
  4. What are match and case, and why are they not in keyword.kwlist?
  5. How would you validate a user-provided name before generating Python source code?

4.7 Identifiers

What Are Identifiers?

What Is It?

Identifiers are the names you give to variables, functions, classes, and modules.

Python stores values in objects and uses identifiers as readable references to those objects. An identifier is not the value itself, and assigning a new value to a name does not change the old object.

What Are the Rules?

  • It must start with a letter or underscore (not a digit)
  • It may contain letters, digits, and underscores after that
  • Identifiers are case-sensitive
  • It cannot be a keyword

Naming Conventions

See 4.16 Coding Standards (PEP 8) for the full naming table — snake_case for variables/functions, PascalCase for classes, UPPER_SNAKE_CASE for constants.

Identifier kindRecommended styleExample
Variable or functionsnake_caseretry_count, load_config()
ClassPascalCaseDeploymentConfig
ConstantUPPER_SNAKE_CASEDEFAULT_TIMEOUT
Internal implementation detailOne leading underscore_parse_labels()
Module or packageShort lowercase nameinventory.py

A leading underscore is a naming convention, not a security boundary. It signals that an identifier is intended for internal use. Python’s name-mangling behavior for __name inside a class is separate from this convention.

Valid vs Invalid Identifiers

How Is It Used?

>>> "valid_name".isidentifier()
True
>>> "2invalid".isidentifier()      # cannot start with a digit
False
>>> "my-var".isidentifier()        # hyphens are not allowed
False

Unicode letters can also be valid in identifiers, but ASCII names are usually easier for teams to search, type, review, and operate across shells and CI systems:

service_name = "api"
π = 3.14159

print(service_name, π)

For production automation, prefer descriptive ASCII identifiers unless the project has a deliberate convention for Unicode names.

How Python Validates an Identifier

Python checks the first character, the remaining characters, and whether the name is reserved. A valid-looking name can still fail the final check if it is a keyword.

flowchart TD N[Candidate name] --> F{isidentifier?} F -->|No| I[Invalid syntax or characters] F -->|Yes| K{iskeyword?} K -->|Yes| R[Reserved by Python] K -->|No| V[Safe identifier shape]
import keyword


def can_be_identifier(candidate: str) -> bool:
	return candidate.isidentifier() and not keyword.iskeyword(candidate)


for candidate in ("pod_name", "9pods", "class", "service-name"):
	print(candidate, can_be_identifier(candidate))

Expected output:

pod_name True
9pods False
class False
service-name False

Reserved Words

The 35 keywords from 4.6 Keywords cannot be used as identifiers, even though they otherwise look like valid names — Python’s parser treats them as fixed syntax rather than user-defined names.

Built-in names are different from keywords. Names such as list, type, and id are valid identifiers, but reusing them hides the built-in and can cause a later failure:

list = ["api", "worker"]
print(list("abc"))  # TypeError: 'list' object is not callable

Use names such as services or service_names instead. A linter can catch many shadowing problems before they reach CI.

Scope and Rebinding

An identifier can refer to different objects in different scopes. A local assignment normally does not change a global name with the same spelling:

environment = "production"


def deploy():
	environment = "staging"
	return environment


print(deploy())       # staging
print(environment)    # production

The function’s environment is a local identifier. Use global or nonlocal only when rebinding an outer name is intentional; passing values and returning results is usually easier to test.

Practical DevOps Example

Clear identifiers make automation scripts easier to review and reduce mistakes when several configuration values have similar meanings:

import os


DEFAULT_REGION = "us-east-1"


def deployment_target() -> tuple[str, str]:
	region = os.getenv("AWS_REGION", DEFAULT_REGION)
	cluster_name = os.environ["EKS_CLUSTER_NAME"]
	return cluster_name, region


cluster_name, region = deployment_target()
print(f"Deploying to {cluster_name} in {region}")

DEFAULT_REGION communicates a constant, while cluster_name and region communicate values that can change. Avoid vague names such as x, data, or value when the script performs infrastructure operations.

Quick Interview Answer

“An identifier must start with a letter or underscore, contain only letters/digits/underscores after that, and can’t be a keyword. Identifiers are case-sensitive — total and Total are different names. Use str.isidentifier() to check programmatically, and keyword.iskeyword() to check whether a valid-looking name is actually reserved.”

Troubleshooting Checklist

When Python reports an invalid identifier:

  1. Check that the first character is not a digit.
  2. Replace hyphens, spaces, and punctuation with underscores.
  3. Check the spelling against keyword.kwlist if the name looks like Python syntax.
  4. Look for invisible or non-ASCII characters copied from documentation.
  5. Check for accidental built-in shadowing such as list = ... or id = ....
  6. Run python -m py_compile your_script.py to catch parsing errors before executing the script.

Interview Follow-up Questions

  1. What is the difference between an identifier and a variable in Python?
  2. Why are total and Total different identifiers?
  3. How do str.isidentifier() and keyword.iskeyword() differ?
  4. Is _name private in Python?
  5. What problem is caused by assigning a value to list or id?
  6. How does local scope affect an identifier with the same name as a global variable?

Common Mistakes

  • Starting an identifier with a digit (2invalid) — always a SyntaxError.
  • Using a hyphen instead of an underscore (my-var) — Python parses the hyphen as a minus operator, not part of the name.
  • Assuming identifiers are case-insensitive, then being confused why Username and username are treated as two separate variables.
  • Confusing a valid identifier with a good name — data may be legal, but pod_metadata is more useful to the next engineer.
  • Treating a leading underscore as access control — it communicates intent but does not prevent access from other code.

4.8 Variables

Declaring Variables

What Is It?

Unlike many languages, Python has no separate declaration step — a variable comes into existence the moment you assign a value to it.

Why Is It Used?

Less boilerplate, faster to write.

How Is It Used?

Just write name = value.

>>> username = "deploy_bot"
>>> username
'deploy_bot'

Python variables are names, not fixed storage boxes with permanent types. The name username refers to a string object, and a later assignment can make it refer to a different object.

Assignment

The = operator binds a name to a value (technically, to an object in memory). Re-assigning simply points the name at a new object; it doesn’t modify the old one.

>>> retries = 3
>>> retries = retries + 1
>>> retries
4

Names and Objects

Assignment creates or updates a binding between a name and an object. It does not copy the object automatically.

flowchart LR A[retries] --> O1[(integer object 3)] B[retries = retries + 1] --> O2[(integer object 4)] A -. name is rebound .-> O2

The old integer object is not changed. The name is simply rebound to the result of retries + 1.

Use == to compare values and is only to compare object identity. In normal application code, is is most commonly used with the singleton None:

value = None

if value is None:
	print("No value was supplied")

Multiple Assignment

What Is It?

Assigning several variables in one statement, either to different values or the same one.

Why Is It Used?

It reduces repetition when initializing related variables together.

>>> a, b, c = 1, 2, 3    # unpack three values at once
>>> a, b, c
(1, 2, 3)

>>> x = y = z = 0        # all three names point at the same value
>>> x, y, z
(0, 0, 0)

Python also supports starred unpacking when the number of remaining values is variable:

>>> first, *middle, last = [10, 20, 30, 40]
>>> first, middle, last
(10, [20, 30], 40)

The number of values must still be compatible with the assignment pattern. Otherwise Python raises ValueError.

Mutable Objects and Aliasing

Multiple names can refer to the same mutable object. Mutating the object is visible through every alias, while rebinding one name is not:

>>> primary = []
>>> secondary = primary
>>> secondary.append("api")
>>> primary
['api']
>>> secondary = ["worker"]
>>> primary
['api']

This is why a mutable default argument is dangerous:

def add_service(service_name, services=[]):
	services.append(service_name)
	return services

The same list is reused between calls. Use None as a sentinel and create the list inside the function instead:

def add_service(service_name, services=None):
	if services is None:
		services = []
	services.append(service_name)
	return services

Type Annotations

Annotations document the expected type without changing Python’s dynamic runtime behavior. Python does not enforce the annotation by itself:

retries: int = 3
service_name: str = "api"

retries = "three"  # Allowed at runtime, but a type checker should report it.

Use annotations for function boundaries and configuration data, then use a type checker such as Pyright or mypy in CI when the project requires static checks.

Dynamic Typing

Covered in depth in Introduction to Python, Section 1.3 — a variable’s type is simply whatever its current value’s type is, and can change on reassignment:

>>> v = 10
>>> type(v)
<class 'int'>
>>> v = "text"
>>> type(v)
<class 'str'>

Dynamic typing does not mean a value has no type. Every object has a type; it means the name is not permanently restricted to one type by the language.

Variables in DevOps Scripts

Keep external configuration at the boundary of a script, validate it, and pass explicit values to the functions that need them:

import os


def get_deployment_config() -> tuple[str, str]:
	cluster_name = os.environ.get("EKS_CLUSTER_NAME")
	region = os.environ.get("AWS_REGION", "us-east-1")

	if not cluster_name:
		raise RuntimeError("EKS_CLUSTER_NAME is required")

	return cluster_name, region


cluster_name, region = get_deployment_config()
print(f"Deploying {cluster_name} in {region}")

This keeps secrets and environment-specific values outside the source code while making the function’s inputs clear and testable. Do not print secret environment variables in CI logs.

Variable Scope

A name created inside a function is local to that function unless the code explicitly refers to an outer scope:

environment = "production"


def target_environment():
	environment = "staging"
	return environment


print(target_environment())  # staging
print(environment)           # production

Prefer function parameters and return values over global state. Explicit data flow is easier to test and safer when automation runs concurrently.

Quick Interview Answer

“A Python variable is just a name bound to an object — there’s no separate declaration step, no type annotation required. x = y = z = 0 binds all three names to the same object; a, b, c = 1, 2, 3 unpacks three values in one line. Because a name is just a label, reassigning it to a different type is perfectly legal — that’s what ‘dynamically typed’ means.”

Troubleshooting Checklist

When a variable behaves unexpectedly:

  1. Print or inspect its value and type(value).
  2. Check whether the name was accidentally rebound later in the function.
  3. For mutable values, inspect whether two names refer to the same object with is.
  4. Use == for value comparison and is None for the absence of a value.
  5. Check unpacking counts when Python raises ValueError.
  6. Review environment-variable defaults and validate required configuration before use.

Interview Follow-up Questions

  1. Is a Python variable a box, a pointer, or a name bound to an object?
  2. What is the difference between rebinding and mutating an object?
  3. Why does x = y = [] create an aliasing problem?
  4. What is the difference between == and is?
  5. Do type annotations enforce types at runtime?
  6. Why is None a safer default sentinel than an empty mutable list?

Common Mistakes

  • Assuming x = y = [] creates two separate lists — it creates one list object that both names point to, so mutating it through either name affects both.
  • Forgetting that reassignment doesn’t mutate the old value — it just points the name at a new object, leaving the old one unchanged (and eventually garbage-collected if nothing else references it).
  • Using a mutable list or dictionary as a default function argument — the same object persists across calls.
  • Using is for ordinary value comparison — use == unless object identity is specifically what you need.
  • Believing annotations enforce types — they document intent and support tools, but Python does not enforce them automatically.

4.9 Constants

Concept of Constants

What Is It?

A value that’s meant to never change after it’s set.

Why Doesn’t Python Enforce It?

Python has no true language-enforced constant — unlike const in other languages, nothing stops you from reassigning an UPPER_CASE name. “Constants” in Python are a naming convention, not a compiler guarantee.

Naming Convention

How Is It Used?

Constants are written in UPPER_SNAKE_CASE to visually signal “don’t reassign this” to anyone reading the code:

MAX_RETRIES = 3
DEFAULT_TIMEOUT = 30
API_BASE_URL = "https://api.example.com"

Best Practices

  • Define constants at the top of the module, right after imports
  • Use UPPER_SNAKE_CASE consistently so violations of the convention stand out
  • For values that truly must be immutable, use a tuple or Enum rather than relying on naming alone

Quick Interview Answer

“Python has no const keyword — a ‘constant’ is purely a naming convention: UPPER_SNAKE_CASE, defined at the top of a module, meant to signal ‘don’t reassign this.’ Nothing at the language level actually prevents reassignment. If true immutability matters, use a tuple or an Enum instead of relying on the naming convention alone.”

Common Mistakes

  • Believing UPPER_CASE = value is protected from reassignment the way const is in JavaScript or Java — it is not; it’s purely convention.
  • Defining constants scattered throughout a module instead of grouped at the top, making them harder for a reader to find.

4.10 Literals

A literal is a value written directly in source code, as opposed to one computed at runtime — 42, "hello", and True are all literals.

Numeric Literals

>>> type(10)         # int
<class 'int'>
>>> type(10.5)        # float
<class 'float'>
>>> type(1 + 2j)      # complex
<class 'complex'>

String Literals

Covered in depth in Chapter 9: Strings — single, double, and triple quotes all produce str literals.

>>> type("hello")
<class 'str'>

Boolean Literals

>>> type(True)
<class 'bool'>

None

What Is It?

Python’s explicit “no value” placeholder — its own type, distinct from 0, False, or an empty string.

Why Is It Used?

To represent the deliberate absence of a value, e.g. a function that has nothing meaningful to return.

>>> type(None)
<class 'NoneType'>

Collection Literals

>>> type([1, 2])        # list
<class 'list'>
>>> type((1, 2))         # tuple
<class 'tuple'>
>>> type({1, 2})         # set
<class 'set'>
>>> type({"a": 1})       # dict
<class 'dict'>

Quick Interview Answer

“A literal is a value written directly into source code rather than computed — numeric (int, float, complex), string, boolean (True/False), None, and the collection literals [], (), {} for list/tuple/set/dict. None is its own distinct type, NoneType — not the same as 0, False, or "".”

Common Mistakes

  • Treating None, 0, False, and "" as interchangeable — they’re all “falsy” in a boolean context, but None specifically represents the absence of a value, not a zero or empty one.
  • Writing {} expecting an empty set — {} is actually an empty dict; an empty set requires set().

4.11 Input Statement

Python Input

What Is input()?

input() reads one line from standard input, optionally displays a prompt, waits until the user presses Enter, removes the line ending, and returns the remaining text as a string.

It is useful for small interactive command-line scripts. It is not a replacement for a full command-line argument parser, environment variables, or a secret manager in production automation.

Basic Syntax

value = input(prompt)

The prompt argument is optional and must be a string. Python writes it to the console without adding a newline, then waits for input.

name = input("Enter your name: ")
print(f"Hello, {name}!")

Example interaction:

Enter your name: Alice
Hello, Alice!

The prompt can be omitted:

print("Enter the environment name:")
environment = input()

The value is assigned only when the user enters a line. input() does not automatically validate, convert, or hide anything typed by the user.

Important Properties

It Always Returns str

Even when the user enters digits, the result is text.

replica_text = input("Number of replicas: ")
print(type(replica_text))  # <class 'str'>

replicas = int(replica_text)
print(replicas + 1)

Without conversion, "3" + "1" produces "31", not 4, and "3" + 1 raises TypeError.

It Reads One Line

The user can type spaces inside the line, but pressing Enter finishes that call. To read several lines, call input() repeatedly or use a loop.

first_line = input("First line: ")
second_line = input("Second line: ")

It Removes the Trailing Newline

Unlike sys.stdin.readline(), the returned value does not include the final \\n. It can still contain leading or trailing spaces, so .strip() is often useful.

service = input("Service name: ").strip()

Common Ways to Use input()

Read Plain Text

region = input("AWS Region: ").strip()

Convert to an Integer

port = int(input("Port: "))

Use this only when invalid text should stop the program. For user-facing tools, catch ValueError and show a helpful message.

Convert to a Decimal Number

threshold = float(input("CPU threshold (%): "))

For currency or exact decimal values, prefer decimal.Decimal instead of float.

Parse a Boolean Safely

Do not use bool(input("Enable feature? ")): every non-empty string, including "no", is truthy.

answer = input("Enable monitoring? (yes/no): ").strip().lower()
if answer in {"yes", "y"}:
  monitoring_enabled = True
elif answer in {"no", "n"}:
  monitoring_enabled = False
else:
  raise ValueError("Enter yes or no")

Read Several Values from One Line

split() separates whitespace-delimited values. Convert each item when numeric input is expected.

raw_ports = input("Ports, separated by spaces: ")
ports = [int(port) for port in raw_ports.split()]
print(ports)

For comma-separated values, provide the separator explicitly:

raw_tags = input("Tags, separated by commas: ")
tags = [tag.strip() for tag in raw_tags.split(",") if tag.strip()]

Use a Default Value

An empty response can mean “use the default”. The or expression handles that case.

region = input("AWS Region [us-east-1]: ").strip() or "us-east-1"

Ask Until the Value Is Valid

This is the standard pattern for a small interactive tool.

while True:
  try:
    replicas = int(input("Replicas (1-20): "))
    if 1 <= replicas <= 20:
      break
    print("Enter a number from 1 to 20.")
  except ValueError:
    print("Enter a whole number.")

print(f"Deploying {replicas} replicas.")

Choose from a Menu

actions = {"1": "status", "2": "logs", "3": "restart"}
choice = input("1) status  2) logs  3) restart: ").strip()

if choice not in actions:
  raise ValueError("Unknown action")
print(f"Selected action: {actions[choice]}")

Read Multiple Lines

commands = []
for number in range(3):
  commands.append(input(f"Command {number + 1}: ").strip())

For an unknown number of lines, use a sentinel such as "done":

lines = []
while True:
  line = input("Enter a line, or done: ").strip()
  if line.lower() == "done":
    break
  lines.append(line)

Handling Errors and Interruptions

Invalid conversion raises ValueError. End-of-file, common when stdin is redirected or a CI job provides no input, raises EOFError. Ctrl+C raises KeyboardInterrupt.

try:
  environment = input("Environment: ").strip()
  replicas = int(input("Replicas: "))
except ValueError:
  print("Replicas must be a whole number.")
except EOFError:
  print("No interactive input was provided.")
except KeyboardInterrupt:
  print("\nOperation cancelled.")

input() and Standard Input

input() reads from standard input, so a shell can pipe a value into a script. This is useful for simple, non-secret values.

echo "staging" | python deploy.py
environment = input().strip()
print(f"Received environment: {environment}")

For larger streams, use sys.stdin instead of calling input() one line at a time:

import sys

for line in sys.stdin:
  print(f"Processing: {line.strip()}")

Do not pipe passwords, tokens, or cloud credentials through shell commands because they can appear in shell history, process output, or CI logs.

Cloud and DevOps Examples

Select an AWS Region and Profile

This example uses input for choices while leaving authentication to the AWS SDK credential chain. The user should already have a configured AWS profile; credentials should never be requested with input().

import boto3

region = input("AWS region [us-east-1]: ").strip() or "us-east-1"
profile = input("AWS profile [default]: ").strip() or "default"

session = boto3.Session(profile_name=profile, region_name=region)
ec2 = session.client("ec2")

response = ec2.describe_instances(
  Filters=[{"Name": "instance-state-name", "Values": ["running"]}]
)
running = sum(
  len(reservation["Instances"])
  for reservation in response["Reservations"]
)
print(f"Running instances in {region}: {running}")

In a non-interactive pipeline, pass the region and profile as command-line arguments or environment variables instead. Interactive input can block a deployment job indefinitely.

Confirm a Potentially Destructive AWS Action

Always require an explicit confirmation before deleting or changing infrastructure.

resource_id = input("EC2 instance ID to stop: ").strip()
confirmation = input(
  f"Type STOP to stop {resource_id}, or anything else to cancel: "
).strip()

if confirmation == "STOP":
  print(f"Would stop {resource_id}")
  # ec2.stop_instances(InstanceIds=[resource_id])
else:
  print("Cancelled.")

The example leaves the API call commented out so it cannot change a real account when copied for learning.

Select a Kubernetes Context and Namespace

The input is validated before it is passed to kubectl. In production automation, prefer a reviewed allowlist and avoid constructing shell commands with shell=True.

import subprocess

context = input("kubectl context: ").strip()
namespace = input("Namespace [default]: ").strip() or "default"

if not context or not namespace:
  raise ValueError("Context and namespace are required")

subprocess.run(
  ["kubectl", "--context", context, "--namespace", namespace, "get", "pods"],
  check=True,
)

Build a Docker Image Tag

Validate a release value before using it in a Docker command or CI job.

import re

image = input("Image name [orders-api]: ").strip() or "orders-api"
tag = input("Image tag [latest]: ").strip() or "latest"

if not re.fullmatch(r"[a-z0-9][a-z0-9_.-]*", image):
  raise ValueError("Invalid image name")
if not re.fullmatch(r"[A-Za-z0-9][A-Za-z0-9_.-]*", tag):
  raise ValueError("Invalid image tag")

print(f"docker build --tag {image}:{tag} .")

Choose a Log Level for a Diagnostic Script

import logging

level_names = {"debug": logging.DEBUG, "info": logging.INFO, "warning": logging.WARNING}
level = input("Log level [info]: ").strip().lower() or "info"

if level not in level_names:
  raise ValueError(f"Choose one of: {', '.join(level_names)}")
logging.basicConfig(level=level_names[level])
logging.info("Diagnostic collection started")

Collect a CI/CD Deployment Choice

This is suitable for a manually started local release helper, not an unattended pipeline.

environments = {"1": "dev", "2": "staging", "3": "production"}
choice = input("Deploy to 1) dev  2) staging  3) production: ").strip()

environment = environments.get(choice)
if environment is None:
  raise ValueError("Invalid deployment environment")

if environment == "production":
  confirm = input("Type DEPLOY to continue: ").strip()
  if confirm != "DEPLOY":
    raise SystemExit("Production deployment cancelled")

print(f"Starting deployment to {environment}")

input() Versus Other Configuration Methods

SituationPreferred method
A person chooses an option in a local scriptinput()
A scheduled job needs a valueEnvironment variable or config file
A CI job needs typed optionsargparse command-line arguments
A password or token is requiredgetpass.getpass() or a secret manager
Many structured settings are requiredYAML, JSON, or a typed configuration model
A script processes piped log datasys.stdin

For a hidden password prompt, use getpass:

from getpass import getpass

password = getpass("Password: ")

Even getpass() is usually the wrong place for production cloud credentials. Use IAM roles, workload identity, a CI secret store, or the cloud provider’s credential chain.

Quick Interview Answer

“input(prompt) writes an optional prompt, waits for one line from standard input, removes the trailing newline, and always returns a string. I use .strip() for whitespace, convert explicitly with int() or float(), validate in a loop, and catch ValueError, EOFError, and KeyboardInterrupt where appropriate. For production DevOps automation, I use argparse, environment variables, or a secret manager when the script must be non-interactive or handle sensitive values.”

Common Mistakes

  • Assuming input() returns a number. Convert and validate it explicitly.
  • Using bool(input(...)) to parse yes/no. Any non-empty string is True.
  • Forgetting .strip(), causing a valid-looking choice with extra spaces to fail validation.
  • Asking for passwords, API keys, or cloud credentials with input(), which displays the typed value.
  • Using input() in unattended CI/CD jobs, where no person is available to answer and the job can hang.
  • Passing unchecked input into a shell command. Prefer an argument list with subprocess.run() and validate allowed values.
  • Treating an empty response as a valid resource name, region, namespace, or deployment target.

4.12 Output Statement

Python Output

What Is print()?

print() writes values to a text stream, usually the terminal. It accepts zero or more positional values, converts them to text, separates them with sep, and appends end.

print("Hello", "World")

Output:

Hello World

print() is useful for human-readable status messages, command-line tools, quick debugging, and small scripts. For long-running production services, use the logging module so messages have levels, timestamps, and configurable destinations.

Basic Syntax

print(*objects, sep=" ", end="\n", file=None, flush=False)
ArgumentPurposeDefault
*objectsValues to write; any number is allowedNone
sepText placed between valuesOne space
endText written after the final valueNewline (\\n)
fileText stream receiving the outputStandard output (sys.stdout)
flushForce buffered output to be written immediatelyFalse

All arguments after the values are keyword arguments. This is valid:

print("deployment", "complete", sep=" | ", end="!\n")

Output:

deployment | complete!

Printing Values

Print One Value

service = "orders-api"
print(service)

Print Multiple Values

Python converts each value to text for display.

service = "orders-api"
replicas = 3
healthy = True

print(service, replicas, healthy)

Output:

orders-api 3 True

This is usually clearer than manually concatenating strings, and it avoids type errors such as trying to concatenate a string and an integer.

Print No Value

Calling print() with no arguments writes one blank line.

print("Deployment summary")
print()
print("Status: healthy")

Print Collections

Lists and dictionaries are displayed using their normal Python representation.

regions = ["us-east-1", "eu-west-1"]
configuration = {"replicas": 3, "environment": "staging"}

print(regions)
print(configuration)

For machine-readable output, use JSON rather than relying on the Python representation. See Structured JSON Output.

The sep Parameter

sep controls the text inserted between multiple positional values. Its default is one space.

print("api", "staging", "healthy", sep=" | ")

Output:

api | staging | healthy

Useful separators include commas, colons, tabs, and empty text:

print("2026-09-23", "10:30:00", "INFO", sep=" ")
print("name", "status", "latency_ms", sep=",")
print("loading", end="")
print(".", end="")
print(".")

sep is ignored when there is only one object.

The end Parameter

end controls what is written after the final value. The default is a newline.

print("first")
print("second")

Output:

first
second

Set end="" to continue on the same line:

print("Checking", end=" ")
print("database", end=" ")
print("OK")

Output:

Checking database OK

It can also be used for a simple progress display:

for step in range(1, 4):
  print(f"step {step}", end=" ", flush=True)
print("complete")

For a real progress bar, use a library such as tqdm rather than building terminal control behavior manually.

Escape Characters

Escape characters represent special characters inside strings. The complete reference is 4.13 Escape Characters.

print("Line 1\nLine 2")
print("Name\tStatus")
print("She said, \"healthy\"")

Output:

Line 1
Line 2
Name    Status
She said, "healthy"

Use a raw string for text containing many backslashes, such as a Windows path or regular expression:

print(r"C:\\Users\\deploy\\logs")

Formatted Output with f-Strings

f-strings are the preferred way to place expressions inside output text.

service = "payments-api"
status = "healthy"
latency_ms = 42.7

print(f"{service}: {status}, latency={latency_ms:.1f} ms")

Common formatting options:

completion = 0.9375
replicas = 3

print(f"Completion: {completion:.1%}")
print(f"Replicas: {replicas:02d}")
print(f"Cost: ${12.5:,.2f}")

Output:

Completion: 93.8%
Replicas: 03
Cost: $12.50

Expressions can be placed inside braces, but keep complicated logic outside the output statement:

is_healthy = replicas > 0
print(f"Ready: {is_healthy}")

str() and repr() Output

print() uses the readable string form of values. Use repr() when debugging and you need to see quotes, escape characters, or whitespace.

value = "  staging\n"
print(value)
print(repr(value))

Output:

  staging

'  staging\n'

This distinction is useful when diagnosing configuration values that contain unexpected spaces or newlines.

Writing Output to a File or Stream

By default, output goes to sys.stdout. The file parameter sends it to another text stream.

with open("deployment-report.txt", "w", encoding="utf-8") as report:
  print("Deployment completed", file=report)
  print("Environment: staging", file=report)

Use sys.stderr for errors or diagnostic messages that should be separate from normal output:

import sys

print("Deployment started")
print("Warning: replica count is low", file=sys.stderr)

This separation allows a shell or CI system to redirect normal results and errors independently.

Buffering and flush=True

Some streams buffer output before writing it. This can make a long-running script appear stuck, especially when output is sent to a pipe or CI log. flush=True requests immediate flushing.

import time

for item in ["download", "extract", "deploy"]:
  print(f"Starting {item}...", flush=True)
  time.sleep(1)

Use flushing for progress or heartbeat messages, not for every normal log line without a reason.

Structured JSON Output

Human-readable output is useful at a terminal, but automation works better with stable JSON keys.

import json

result = {
  "service": "orders-api",
  "environment": "staging",
  "healthy": True,
  "replicas": 3,
}

print(json.dumps(result))

Pretty-print JSON for people reviewing a report:

print(json.dumps(result, indent=2, sort_keys=True))

Do not print secrets, access tokens, passwords, or complete cloud API responses if they may contain sensitive data.

print() Versus logging

Use print() for short scripts, simple user feedback, and quick demonstrations. Use logging for production services and automation that needs levels, timestamps, handlers, or filtering.

import logging

logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
logging.info("Deployment started")
logging.warning("Rollback threshold is near")

Avoid logging credentials, authorization headers, full environment variables, or sensitive customer data.

Cloud and DevOps Examples

Display an AWS Resource Summary

The SDK response is converted into a small stable summary rather than printing the entire response, which may contain unnecessary details.

import boto3

ec2 = boto3.client("ec2", region_name="us-east-1")
response = ec2.describe_instances(
  Filters=[{"Name": "instance-state-name", "Values": ["running"]}]
)

running_instances = [
  instance
  for reservation in response["Reservations"]
  for instance in reservation["Instances"]
]

print(f"Running instances: {len(running_instances)}")
for instance in running_instances:
  print(instance["InstanceId"], instance.get("InstanceType"), sep=" | ")

In real automation, obtain the region and credentials from the AWS configuration or workload identity rather than hardcoding secrets in output code.

Print an AWS Deployment Result as JSON

import json

deployment = {
  "application": "orders-api",
  "region": "us-east-1",
  "version": "2026.09.23",
  "status": "success",
}
print(json.dumps(deployment))

This format can be consumed by a CI job, monitoring system, or another script.

Report Kubernetes Pod Status

import subprocess

result = subprocess.run(
  ["kubectl", "get", "pods", "--namespace", "staging", "--no-headers"],
  capture_output=True,
  text=True,
  check=True,
)

for line in result.stdout.splitlines():
  columns = line.split()
  if len(columns) >= 3:
    print(f"{columns[0]} | ready={columns[1]} | status={columns[2]}")

Using an argument list avoids shell parsing problems. Do not pass unchecked user input into shell=True.

Show Docker Image Build Progress

import subprocess

image_tag = "orders-api:2026.09.23"
print(f"Building {image_tag}...", flush=True)
subprocess.run(["docker", "build", "--tag", image_tag, "."], check=True)
print(f"Built {image_tag}")

For detailed command output, let the subprocess inherit the terminal or CI stream instead of capturing and printing it only after the command finishes.

Produce a CI/CD Check Result

import json

checks = {
  "unit_tests": "passed",
  "security_scan": "passed",
  "image_scan": "warning",
}

failed = [name for name, status in checks.items() if status == "failed"]
summary = {"checks": checks, "failed": failed, "success": not failed}
print(json.dumps(summary))

A CI job can parse the JSON and decide whether a pipeline stage should continue. Human-facing status messages can still use f-strings.

Print a Log Scanner Summary

from collections import Counter

levels = ["INFO", "ERROR", "INFO", "WARNING", "ERROR"]
counts = Counter(levels)

for level, count in sorted(counts.items()):
  print(f"{level:<8} {count:>4}")

Output:

ERROR       2
INFO        2
WARNING     1

Output Design for Automation

Choose output based on who or what consumes it:

ConsumerGood output choice
Person at a terminalClear f-string messages with useful spacing
Shell pipelineOne stable value per line or a documented delimiter
CI/CD jobJSON or a clearly defined exit status and stream
Monitoring systemStructured logs through logging
Error handlingsys.stderr plus a non-zero process exit code
Large command outputStream it rather than storing the entire result

Output should be stable when another program parses it. Avoid adding decorative text to a machine-readable stream such as JSON.

Quick Interview Answer

“print() writes values to a text stream. It accepts any number of objects and supports sep between objects, end after the final object, file for a destination such as a file or sys.stderr, and flush for immediate output. I use f-strings for human-readable messages, json.dumps() for machine-readable results, and logging instead of scattered print() calls in production services.”

Common Mistakes

  • Forgetting that print() adds a newline by default and then getting unexpected line breaks.
  • Using string concatenation with numbers instead of multiple arguments or an f-string.
  • Printing an entire cloud SDK response, which can produce noisy output or expose sensitive fields.
  • Sending errors to standard output instead of sys.stderr, making shell and CI redirection confusing.
  • Using human-readable labels in output that another program expects to parse as JSON.
  • Omitting flush=True when a long-running progress or heartbeat message must appear immediately.
  • Using print() as the logging system for a production service instead of configuring logging.
  • Logging passwords, tokens, authorization headers, or complete environment variables.

4.13 Escape Characters

Escape sequences let you embed special or hard-to-type characters inside an ordinary quoted string, using a backslash followed by a code.

EscapeMeaningExample
\nNewline"a\nb" → two lines
\tTab"a\tb" → "a b"
\\Literal backslash"a\\b" → a\b
\'Literal single quote'it\'s' → it's
\"Literal double quote"say \"hi\"" → say "hi"
\rCarriage returnused in some line-ending formats
\bBackspacemoves cursor back one position
>>> print("a\nb")
a
b
>>> print("a\tb")
a    b
>>> print("say \"hi\"")
say "hi"

The full escape-character reference, including octal/hex codes, is covered in 9.3 Escape Characters.

Quick Interview Answer

“An escape sequence is a backslash followed by a code that represents a special character inside a string — \n for newline, \t for tab, \\ for a literal backslash, \'/\" for a quote that would otherwise end the string early.”

Common Mistakes

  • Forgetting to escape a quote character that matches the string’s own delimiter, causing the string to end early and raise a SyntaxError.
  • Confusing \n (a two-character escape sequence Python turns into one newline byte) with an actual line break typed into a triple-quoted string — both work, but they’re not the same mechanism.

4.14 Code Blocks

A code block is any indented group of statements introduced by a colon-terminated header line. All four kinds below follow the exact same indentation rule from 4.4 Indentation.

flowchart TD CB["Code Block\n(colon-terminated header\n+ indented body)"] CB --> IF["if Block"] CB --> LOOP["Loop Block"] CB --> FUNC["Function Block"] CB --> CLASS["Class Block"]

if Blocks

if status == "active":
    print("Service is running")

Loop Blocks

for server in ["web01", "web02"]:
    print(f"Checking {server}")

Function Blocks

def is_valid(code):
    return 200 <= code < 300

Class Blocks

class Server:
    def __init__(self, name):
        self.name = name

Quick Interview Answer

“A code block is any colon-terminated header line followed by an indented body — if, for/while, def, and class all use the exact same mechanism. There’s no separate ‘block’ keyword in Python; indentation alone marks where a block starts and ends.”

Common Mistakes

  • Assuming each block type has different indentation rules — they don’t; if, loops, functions, and classes all follow the identical colon + indent pattern.
  • Forgetting a return inside a function block and getting None back instead of the expected value.

4.15 Multiple Statements & Line Continuation

Three ways Python’s line-based syntax can be bent — combining lines together, or splitting one statement across several lines.

Semicolon

What Is It?

Separates multiple simple statements placed on one physical line.

Why Is It Rarely Used?

Rarely used by convention (see 4.16 Coding Standards (PEP 8)), but valid syntax.

>>> a = 1; b = 2; print(a + b)
3

Backslash

What Is It?

An explicit line-continuation marker: the backslash at the end of a line tells Python “this statement isn’t finished yet, keep reading the next line.”

total = 1 + \
        2 + \
        3
>>> total
6

Implicit Continuation

What Is It?

Inside any open bracket — (), [], or {} — Python automatically treats newlines as continuations, no backslash needed.

Why Is It Preferred?

It’s visually cleaner and less error-prone than backslashes (a trailing space after a backslash silently breaks it).

nums = (1 +
        2 +
        3)
>>> nums
6

Quick Interview Answer

“Python offers three ways to bend its one-statement-per-line rule: semicolons combine statements onto one line (rarely used), a trailing backslash explicitly continues a statement onto the next line, and any open bracket — (), [], {} — implicitly continues across lines with no backslash needed. Implicit continuation inside brackets is the preferred style since a stray trailing space silently breaks a backslash continuation.”

Common Mistakes

  • Leaving a trailing space after a line-continuation backslash — it silently breaks the continuation and raises a SyntaxError on the next line.
  • Reaching for backslash continuation when the expression is already inside brackets, where implicit continuation would work without the backslash at all.

4.16 Coding Standards (PEP 8)

PEP 8 is Python’s official style guide. This section covers it from the syntax angle — the specific formatting rules that affect how code is laid out on the page.

Indentation

4 spaces per indentation level; never mix tabs and spaces (mixing raises a TabError in Python 3).

Line Length

Limit lines to 79 characters (or up to 99 by some team conventions) — keeps code readable side-by-side in diffs and split editor panes.

Naming

snake_case for variables/functions, PascalCase for classes, UPPER_SNAKE_CASE for constants — the same table introduced in Introduction to Python.

Imports

One import per line, grouped in order: standard library, then third-party packages, then local application imports — with a blank line between each group.

import os
import sys

import boto3

from myapp.utils import helper

Whitespace

Use a single space around operators and after commas; avoid extra spaces right inside brackets or right before a function call’s parentheses.

# PEP 8 compliant
total = price * quantity
func(a, b, c)

# Not PEP 8 compliant
total = price*quantity
func( a,b,c )

Quick Interview Answer

“PEP 8 is Python’s official style guide — 4-space indentation, a 79-character line limit, snake_case/PascalCase/UPPER_SNAKE_CASE naming by identifier type, imports grouped stdlib → third-party → local with one per line, and consistent spacing around operators. Tools like black or flake8 enforce most of it automatically, so it’s rarely a manual judgment call on a real team.”

Common Mistakes

  • Mixing import groups together instead of ordering standard library, then third-party, then local application imports.
  • Inconsistent spacing around operators (price*quantity vs price * quantity) that a formatter like black would catch automatically.
  • Treating PEP 8 as optional style preference rather than the shared convention that keeps a team’s codebase consistently readable.

4.17 Common Syntax Errors

Five error types that account for the vast majority of syntax mistakes, especially for beginners — knowing what each one means makes them far faster to fix.

IndentationError

What Does It Mean?

A block is expected but the indentation is missing or inconsistent.

>>> if True:
... print("bad")
IndentationError: expected an indented block after 'if' statement on line 1

SyntaxError

What Does It Mean?

The general-purpose error for code that doesn’t match Python’s grammar at all — often a missing colon, unmatched bracket, or invalid character sequence.

>>> if True
  File "<stdin>", line 1
    if True
           ^
SyntaxError: expected ':'

NameError

What Does It Mean?

Technically a runtime error, not a syntax error — the code IS valid Python, but references a name that was never defined (often a typo).

>>> print(undefined_var)
Traceback (most recent call last):
NameError: name 'undefined_var' is not defined

Missing Colon

Why Does It Happen?

The single most common syntax mistake for beginners coming from other languages — every compound statement header (if/for/while/def/class) must end in :.

>>> def greet(name)
  File "<stdin>", line 1
    def greet(name)
                   ^
SyntaxError: expected ':'

Unmatched Brackets

What Does It Mean?

An opening (, [, or { without its matching close — Python reports exactly which bracket was never closed.

>>> x = [1, 2, 3
  File "<stdin>", line 1
    x = [1, 2, 3
        ^
SyntaxError: '[' was never closed

Quick Interview Answer

“IndentationError means a block’s indentation is missing or inconsistent. SyntaxError is the general grammar violation — most often a missing colon or an unmatched bracket. NameError is different from both: it’s a runtime error, meaning the code is syntactically valid but references a name that was never defined.”

Common Mistakes

  • Confusing NameError (valid syntax, undefined name, caught only at runtime) with SyntaxError (invalid grammar, caught before anything runs).
  • Not reading the caret (^) in the traceback — it points to exactly where the parser gave up, which is usually the fastest way to find the actual problem.

4.18 Real-World DevOps Examples

Python for DevOps

Seeing correct syntax structure in realistic scripts reinforces the rules covered across this chapter better than isolated snippets — here’s how the pieces come together in practice.

Configuration Scripts

A typical config module: constants at the top, dictionary literal for structured settings — direct application of 4.9 Constants and 4.10 Literals.

# config.py
MAX_RETRIES = 3
TIMEOUT_SECONDS = 30

DATABASE = {
    "host": "localhost",
    "port": 5432,
    "name": "prod_db",
}

Automation Scripts

A file-cleanup script showing imports, a function block, and the entry-point guard together — the full layout from 4.2 Structure of a Python Program in action.

import os

def remove_old_logs(folder, days=7):
    for filename in os.listdir(folder):
        print(f"Checking {filename}")

if __name__ == "__main__":
    remove_old_logs("/var/log/app")

Log Processing Script Structure

A minimal but complete log-scanning script — notice the consistent indentation nesting a for loop inside a with block.

def count_errors(log_path):
    error_count = 0
    with open(log_path) as f:
        for line in f:
            if "ERROR" in line:
                error_count += 1
    return error_count

AWS Script Layout

A boto3-based script following the same import → constants → function → entry-point pattern used throughout this chapter.

import boto3

REGION = "us-east-1"

def list_running_instances():
    ec2 = boto3.client("ec2", region_name=REGION)
    return ec2.describe_instances()

if __name__ == "__main__":
    list_running_instances()

Quick Interview Answer

“Real DevOps scripts follow the same layout as any other Python file: imports, then constants (a config dict, an AWS region, retry counts), then function definitions, then an if __name__ == '__main__': guard. Whether it’s cleaning up log files with os, scanning a log for errors with with open(...), or listing EC2 instances with boto3, the underlying syntax rules — indentation, colons, blocks — never change.”

Common Mistakes

  • Hardcoding values like AWS region or file paths directly inside functions instead of pulling them from module-level constants, making the script harder to reconfigure.
  • Forgetting the with statement when opening a file, leaving it to rely on garbage collection to eventually close the file handle.

4.19 Best Practices

A short checklist distilled from every section in this chapter — apply these consistently and most syntax-level code review comments disappear.

Readable Code

Favor clarity over cleverness — code that’s obvious to read is worth more than code that’s marginally shorter but harder to follow.

Meaningful Names

user_count is instantly understandable; uc or x is not. Good identifier names (see 4.7 Identifiers) do a large part of a program’s documentation for free.

Consistent Formatting

Pick PEP 8 (see 4.16 Coding Standards) or your team’s documented variant, and apply it uniformly — tools like black or autopep8 can enforce this automatically so it’s never a manual judgment call.

Comments

Comment intent and reasoning, not mechanics — and keep docstrings (see 4.5 Comments) on every function or module meant to be reused.

Quick Interview Answer

“Readable Python code comes down to four habits: favor clarity over cleverness, use meaningful identifier names instead of single letters, apply a consistent style (PEP 8, enforced by a formatter rather than manually), and comment the why rather than the what. None of these are enforced by the interpreter — they’re team-level discipline that keeps a codebase maintainable.”

Common Mistakes

  • Relying on memory to enforce a style guide instead of an automated formatter (black) or linter (flake8), which drifts inconsistently across a team over time.
  • Optimizing for the shortest possible line of code at the expense of a reader being able to understand it quickly.

4.20 Interview Questions

Frequently Asked Syntax Questions

  • What is the difference between a keyword and an identifier?
  • Why does Python use indentation instead of braces?
  • What is the difference between a SyntaxError and a NameError?
  • Can you reassign a value to a name written in UPPER_CASE? Why or why not?
  • What does input() return, and why does that matter for arithmetic?
  • What is the purpose of if __name__ == "__main__":?

Practical Coding Questions

  • Write a function with a docstring and demonstrate retrieving it with help().
  • Fix this snippet’s IndentationError:
    if True:
    print("hi")
    
  • Show two different ways to continue a long expression across multiple lines.
  • Given a list of numeric strings from input(), write code that sums them as integers.
  • Identify and fix the missing colon in a given faulty function definition.

Quick Interview Answer

“The syntax topics that come up most in interviews are: why Python uses indentation instead of braces (it’s the block-delimiting syntax, not a style choice), the difference between a keyword and an identifier, why SyntaxError is caught before any code runs while NameError is a runtime error, why input() always returns a str, and what if __name__ == '__main__': actually does — letting a file work both as a script and as an importable module.”

Common Mistakes

  • Answering “Python doesn’t have syntax errors, just runtime errors” — conflating the two error categories is one of the fastest ways to lose credibility on a syntax question.
  • Not being able to explain why if __name__ == "__main__": matters, only that “you’re supposed to write it” — interviewers usually follow up asking what breaks without it.

Syntax: Chapter Practice

Mini lab

Fix the indentation in a loop that prints two server names.

Expected output: web-01 and db-01 on separate lines.

HintUse the chapter examples to choose the operation, then run your version before opening the solution.
Show solution
for name in ["web-01", "db-01"]:
    print(name)

Knowledge check

What defines a block in Python?

Check your answerIndentation, introduced after a statement ending in a colon.

Finish the chapter

  • Explain each operation without reading the lesson.
  • Change the sample input and predict the output before running it.
  • Revisit the chapter if the result differs from your prediction.

Back to this chapter · Choose the next chapter

5.1 Introduction to Data Types

Python Data Types
flowchart TD PDT["Python Data Types"] PDT --> NUM["Numeric\nint / float / complex / bool"] PDT --> SEQ["Sequence\nstr / list / tuple / range"] PDT --> MAP["Mapping\ndict"] PDT --> SET["Set\nset / frozenset"] PDT --> BIN["Binary\nbytes / bytearray / memoryview"] PDT --> NON["None\nNoneType"]

The built-in data types covered in this chapter, grouped by category.

What Are Data Types?

What Is It?

A data type classifies what kind of value something is (a number, text, a collection) and, in turn, what operations are valid on it.

Why Does It Matter?

The type determines behavior — you can add two ints, but adding an int to a list raises a TypeError.

How Is It Used?

Every value in Python has exactly one type at any given moment, discoverable with type().

>>> type(42)
<class 'int'>
>>> type("hello")
<class 'str'>

Why Data Types Matter

Choosing the right type affects correctness (can this value be negative? can it have duplicates?), performance (list vs set membership testing), and memory usage. A large part of writing good Python is picking the type that matches your data’s real-world shape.

Dynamic Typing

What Is It?

A variable has no fixed type — its type is simply whatever value it currently holds, and reassignment can change that type freely. This was introduced in 4.8 Variables and is worth re-grounding here since it’s foundational to how types work in Python.

>>> x = 10
>>> type(x)
<class 'int'>
>>> x = "now a string"    # same name, completely different type -- perfectly legal
>>> type(x)
<class 'str'>

Real-World Examples

  • A user’s age → int
  • A server’s hostname → str
  • A list of IP addresses to block → list or set
  • An API response body → dict (parsed from JSON)
  • A file’s raw contents → bytes

Quick Interview Answer

“A data type classifies what kind of value something is and what operations are valid on it — Python discovers a value’s type at runtime with type(), since it’s dynamically typed: a variable’s type is whatever value it currently holds, and reassigning it to a different type is perfectly legal. Picking the right type up front affects correctness, performance, and memory usage.”

Common Mistakes

  • Assuming a variable’s type is fixed once assigned — reassignment to a different type is silent and completely legal in Python.
  • Not checking type() or isinstance() before an operation, then getting a TypeError at runtime instead of catching a bad assumption early.

5.2 Python Object Model

Everything Is an Object

What Is It?

In Python, literally everything — numbers, strings, functions, even classes themselves — is an object with its own identity, type, and value.

Why Does It Matter?

This uniformity is why the same tools (type(), dir(), hasattr()) work on any value, no matter what kind it is.

flowchart LR X["x = 5"] --> ID["Identity — id(x)\na unique integer\n(the object's memory address)"] X --> TY["Type — type(x)\n<class 'int'>\nwhat kind of object it is"] X --> VA["Value — x itself\n5\nthe actual data held"]

Every Python value is an object with three properties: identity, type, and value.

Identity

What Is It?

A unique integer identifying an object for its lifetime — effectively its memory address in CPython.

How Is It Used?

id(x) returns it; the is operator compares two objects’ identities.

>>> a = 5
>>> id(a)
11755816

Type

What kind of object it is, and therefore what operations are valid on it. Retrieved with type().

>>> type(a)
<class 'int'>

Value

The actual data the object holds — 5, "hello", [1, 2, 3]. For mutable objects the value can change over the object’s lifetime; for immutable objects it cannot (see 5.9 Mutable vs Immutable Types).

Quick Interview Answer

“Every Python object has three properties: identity (a unique id, effectively its memory address in CPython, from id()), type (what kind of object it is, from type()), and value (the actual data it holds). This uniform model is why the same tools — type(), dir(), hasattr() — work identically on any value, since everything, even a function or a class, is an object.”

Common Mistakes

  • Confusing id() (identity, a memory address) with type() (what kind of object it is) — they answer completely different questions.
  • Comparing objects with is when == (value equality) is what’s actually intended — is checks identity, not value.

5.3 Numeric Data Types

int

Whole numbers of arbitrary precision (Python ints don’t overflow like fixed-width integers in C — they grow as large as memory allows).

>>> type(10)
<class 'int'>
>>> 2 ** 100     # no overflow, however large
1267650600228229401496703205376

float

Decimal (floating-point) numbers, stored using the IEEE 754 double-precision format — the same trade-offs (like imprecise decimal representation) apply as in most other languages.

>>> type(10.5)
<class 'float'>
>>> 0.1 + 0.2     # classic floating-point precision surprise
0.30000000000000004

complex

Numbers with a real and imaginary part, written with a j suffix — used in scientific/engineering computation, rarely in typical DevOps scripting.

>>> type(2 + 3j)
<class 'complex'>

bool

What Is It?

True/False, but technically a SUBCLASS of int (True == 1, False == 0).

Why Does It Matter?

bool values can be used directly in arithmetic, and isinstance(True, int) is True — a common interview gotcha.

>>> type(True)
<class 'bool'>
>>> int(True), int(False)
(1, 0)
>>> isinstance(True, int)    # bool IS an int subclass
True

Type Conversion

How Is It Used?

Explicit conversion between numeric types uses the type’s own name as a function:

>>> int("42")
42
>>> float("3.14")
3.14
>>> str(42)
'42'

Arithmetic Examples

>>> 10 / 3      # true division -- always returns a float
3.3333333333333335
>>> 10 // 3     # floor division -- returns an int-like whole result
3
>>> 10 % 3      # modulo -- the remainder
1
>>> 10 ** 2     # exponentiation
100

Quick Interview Answer

“Python has four numeric types: int (arbitrary precision, no overflow), float (IEEE 754 double-precision, subject to the classic 0.1 + 0.2 != 0.3 imprecision), complex (real + imaginary, j suffix), and bool, which is technically a subclass of int — True == 1, False == 0, and isinstance(True, int) is True. Explicit conversion between them just calls the target type as a function: int("42"), float("3.14").”

Common Mistakes

  • Comparing floats with == directly (0.1 + 0.2 == 0.3 is False) instead of accounting for floating-point imprecision.
  • Forgetting bool is an int subclass, then being surprised that True + True == 2 or that a type-check with isinstance(x, int) also matches booleans.
  • Using / when floor division // was actually intended, or vice versa.

5.4 Sequence Data Types

A sequence is an ordered collection accessible by integer index. All four types below share this trait but differ in mutability and typical use.

str

An immutable sequence of Unicode characters — covered in exhaustive depth in Chapter 9: Strings.

>>> type("hi")
<class 'str'>

list

What Is It?

A MUTABLE, ordered, resizable sequence — Python’s general-purpose “array”.

Why Is It Used?

It’s the default choice for any ordered collection you’ll add to, remove from, or reorder — covered in exhaustive depth in Chapter 10: Lists.

>>> l = [1, 2, 3]
>>> l.append(4)
>>> l
[1, 2, 3, 4]

tuple

What Is It?

An IMMUTABLE, ordered sequence.

Why Is It Used?

It represents a fixed-size, fixed-content record (like a coordinate pair) and can be used as a dict key or set member, unlike a list — covered in exhaustive depth in Chapter 11: Tuples.

>>> t = (1, 2, 3)
>>> t[0]
1
>>> t[0] = 99
Traceback (most recent call last):
TypeError: 'tuple' object does not support item assignment

range

What Is It?

A memory-efficient, immutable sequence of numbers, generated lazily rather than stored all at once.

Why Is It Used?

Iterating a range(1000000) uses almost no memory, unlike building an actual list of a million numbers.

>>> type(range(5))
<class 'range'>
>>> list(range(5))
[0, 1, 2, 3, 4]

Common Operations

Indexing and len() work identically across all sequence types, since they all implement the same sequence protocol:

>>> s, l, t = "hello", [1, 2, 3], (1, 2, 3)
>>> s[0], l[0], t[0]
('h', 1, 1)
>>> len(s), len(l), len(t)
(5, 3, 3)

Quick Interview Answer

“Python has four sequence types sharing the same indexing/len() protocol: str (immutable text), list (mutable, resizable — the general-purpose array), tuple (immutable, fixed-content — usable as a dict key or set member, unlike a list), and range (a lazy, memory-efficient sequence of numbers that never materializes the full list).”

Common Mistakes

  • Trying to mutate a tuple in place (t[0] = 99) instead of recognizing it’s immutable and building a new tuple.
  • Materializing a huge range() into a list() unnecessarily, losing its lazy memory efficiency for no reason.
  • Using a list as a dict key or set member and hitting TypeError: unhashable type: 'list' — reach for a tuple instead.

5.5 Mapping Data Type

dict

What Is It?

Python’s built-in hash map — a mutable, unordered (technically insertion-ordered since 3.7) collection of key-value pairs.

Why Is It Used?

It’s the natural representation for structured, labeled data — config settings, JSON objects, API responses.

>>> d = {"name": "Alice", "age": 30}
>>> type(d)
<class 'dict'>

Key-Value Pairs

How Is It Used?

Every entry maps a unique, hashable key to a value. Keys can be any immutable type (str, int, tuple); values can be anything.

>>> d["name"]
'Alice'
>>> d.get("missing", "default")    # avoids KeyError for absent keys
'default'

Dictionary Use Cases

  • Parsed JSON / API response bodies
  • Configuration settings (key → value)
  • Counting/grouping (e.g. word → frequency)
  • Fast lookups by a unique identifier (e.g. user ID → user record)

Quick Interview Answer

“dict is Python’s built-in hash map — a mutable collection of key-value pairs, insertion-ordered since Python 3.7. Keys must be hashable (immutable types like str, int, tuple); values can be anything. It’s the natural fit for structured, labeled data — config settings, parsed JSON, API responses — and .get(key, default) is the standard way to look up a key without risking a KeyError.”

Common Mistakes

  • Using d[key] when the key might not exist, raising an unhandled KeyError — use .get() with a default instead.
  • Trying to use a mutable type (like a list) as a dict key, which raises TypeError: unhashable type.
  • Assuming dict ordering was always guaranteed — insertion order is only guaranteed from Python 3.7 onward.

5.6 Set Data Types

set

What Is It?

A mutable, unordered collection of unique, hashable values.

Why Is It Used?

Automatic deduplication and fast (O(1) average) membership testing — far faster than checking in on a list for large collections.

>>> s = {1, 2, 3, 2, 1}    # duplicates are automatically dropped
>>> s
{1, 2, 3}

frozenset

The immutable counterpart to set — same behavior, but can’t be modified after creation, which makes it hashable and therefore usable as a dict key or set member itself.

>>> fs = frozenset([1, 2, 3])
>>> fs.add(4)
Traceback (most recent call last):
AttributeError: 'frozenset' object has no attribute 'add'

Unique Values

The most common practical use of set: deduplicating a list in one line.

>>> ips = ["10.0.0.1", "10.0.0.2", "10.0.0.1"]
>>> unique_ips = set(ips)
>>> unique_ips
{'10.0.0.1', '10.0.0.2'}

Quick Interview Answer

“set is a mutable, unordered collection of unique, hashable values — it deduplicates automatically and gives O(1) average membership testing, versus O(n) for a list. frozenset is its immutable counterpart, which makes it hashable and therefore usable as a dict key or as a member of another set. set(some_list) is the standard one-line way to deduplicate.”

Common Mistakes

  • Trying to add to a frozenset after creation — it has no .add(), since it’s immutable by design.
  • Relying on set ordering — sets are unordered, so iteration order isn’t guaranteed the way it is for a list or (since 3.7) a dict.
  • Putting an unhashable value (like a list) into a set, raising TypeError: unhashable type.

5.7 Binary Data Types

These represent raw bytes rather than text — essential whenever Python touches files, network sockets, or any non-text data.

bytes

An immutable sequence of raw 8-bit values (0–255 each). What you get from reading a file in binary mode, or encoding a str (see 9.7 Common String Methods).

>>> b = bytes([104, 105])
>>> b
b'hi'

bytearray

The mutable counterpart to bytes — lets you modify raw byte data in place, useful when building up a binary buffer incrementally.

>>> ba = bytearray(b"hi")
>>> ba[0] = 72
>>> ba
bytearray(b'Hi')

memoryview

What Is It?

A view onto another object’s underlying memory buffer WITHOUT copying it.

Why Is It Used?

Processing large binary data (e.g. a big file read into bytes) efficiently, avoiding the cost of duplicating it in memory.

>>> mv = memoryview(b"hello")
>>> mv[0]
104
>>> bytes(mv)
b'hello'

Binary Data Use Cases

  • Reading/writing files in binary mode (images, archives, executables)
  • Network protocol implementation (raw socket data)
  • Cryptographic operations (hashing, encryption work on bytes)

Quick Interview Answer

“bytes is an immutable sequence of raw 8-bit values — what you get reading a file in binary mode or encoding a string. bytearray is its mutable counterpart, for building up a binary buffer in place. memoryview gives a zero-copy view onto another object’s memory buffer, which matters for processing large binary data efficiently.”

Common Mistakes

  • Trying to mutate a bytes object directly — it’s immutable; use bytearray when in-place modification is needed.
  • Mixing up str and bytes (e.g. writing a str to a file opened in binary mode) — Python raises a TypeError rather than silently converting.
  • Copying a large bytes object unnecessarily instead of using memoryview to operate on it in place.

5.8 NoneType

None Object

What Is It?

Python’s singleton “no value” object — there is exactly one None in a running program, and its type is NoneType.

Why Is It Used?

To explicitly represent the deliberate absence of a value, distinct from any actual data value like 0 or an empty string.

>>> n = None
>>> type(n)
<class 'NoneType'>
>>> n is None    # always compare to None with 'is', not '=='
True

None vs 0 vs Empty String

What Is the Difference?

All three are “falsy” in a boolean context, but they mean different things. 0 is a valid number, "" is a valid (if empty) string — None means no value was ever set at all. Conflating them is a common source of bugs (e.g. treating a legitimate 0 balance as “missing data”).

>>> bool(None), bool(0), bool("")
(False, False, False)    # all falsy, but NOT equal to each other
>>> None == 0, None == ""
(False, False)

Quick Interview Answer

“None is Python’s singleton for ’no value’ — there’s exactly one None object in a running program, of type NoneType. It’s falsy, but it isn’t 0 or \"\" — those are valid, real values. Always compare to None with is, not ==, since is checks identity against that one singleton object rather than relying on equality logic.”

Common Mistakes

  • Comparing to None with == instead of is — is is both faster and the semantically correct check for a singleton.
  • Treating None, 0, and "" as interchangeable “empty” values — they’re all falsy but represent genuinely different things.
  • Using a mutable default of None incorrectly, or conversely forgetting None is the standard sentinel for “no default was given” (see 5.15 Common Mistakes for the mutable-default-argument bug).

5.9 Mutable vs Immutable Types

Definition

What Is It?

A MUTABLE object’s value can be changed in place after creation (its id stays the same); an IMMUTABLE object’s value can never change — any “modification” actually creates a brand-new object.

Why Does It Matter?

This affects correctness (shared references to mutable objects can surprise you), safety (immutable objects are safe to share across threads), and what can be used as a dict key.

flowchart TD subgraph Immutable["Immutable — t = (1, 2); t = t + (3,)"] T1["t → (1, 2)\nid: 0x1001"] -->|"t = t + (3,)"| T2["t → (1, 2, 3)\nid: 0x2002 — a NEW object"] end subgraph Mutable["Mutable — l = [1, 2]; l.append(3)"] L1["l → [1, 2]\nid: 0x3003"] -->|"l.append(3)"| L2["l → [1, 2, 3]\nid: 0x3003 — the SAME object"] end

Mutating a mutable object keeps its identity; “mutating” an immutable one creates a new object.

Examples

MutableImmutable
listtuple
dictstr
setint, float, bool, complex
bytearraybytes, frozenset

Memory Behavior

How Is It Used?

Mutating a list in place leaves its identity (id) unchanged; reassigning a tuple’s contents is impossible, so any “change” produces a new object with a new id:

>>> l = [1, 2, 3]
>>> id(l)
140420020687424
>>> l.append(4)     # in-place mutation
>>> id(l)           # SAME id -- same object, modified
140420020687424

Comparison Table

PropertyMutableImmutable
Can change after creation?YesNo
Safe to use as dict key?No (unhashable)Yes (if hashable)
Safe to share across threads?Requires careInherently safe
id() after modificationUnchangedN/A — new object created

Quick Interview Answer

“A mutable object’s value can change in place — its id stays the same after modification, like a list.append(). An immutable object’s value can never change — any apparent ‘modification’ actually builds a brand-new object with a new id, like tuple + (3,). This is why only immutable, hashable types can be used as dict keys or set members, and why immutable objects are inherently safer to share across threads.”

Common Mistakes

  • Assuming tuple + (x,) modifies the tuple in place — it creates and returns an entirely new tuple; the original is untouched.
  • Sharing a mutable default or a mutable object across function calls without realizing all references point to the same underlying object.
  • Trying to use a list or dict as a dict key, hitting TypeError: unhashable type, instead of reaching for the immutable equivalent (tuple, frozen structure).

5.10 Type Checking

type()

Returns an object’s exact type. Good for debugging/inspection, but generally NOT recommended for validation logic (see isinstance() below), since it does an exact match and ignores subclassing.

>>> type(5) == int
True

isinstance()

What Is It?

Checks whether an object IS an instance of a type, INCLUDING subclasses.

Why Is It Preferred?

It’s preferred over type() for validation because it correctly handles inheritance — e.g. bool is a subclass of int, so isinstance(True, int) is True even though type(True) == int is False.

>>> isinstance(5, int)
True
>>> isinstance(True, int)    # bool IS-A int (subclass)
True
>>> type(True) == int        # exact type match fails here
False

id()

Returns an object’s unique identity (see 5.2 Python Object Model) — used far more often for understanding reference/aliasing behavior than for everyday type checking.

>>> id(5)
11755816

Quick Interview Answer

“type() returns an object’s exact type and does an exact match, which breaks for subclasses. isinstance() checks whether an object IS an instance of a type INCLUDING subclasses, which is why it’s the preferred choice for validation — isinstance(True, int) is True since bool subclasses int, even though type(True) == int is False. id() is a different tool entirely — it returns identity, used for understanding references, not for type checks.”

Common Mistakes

  • Using type(x) == SomeClass for validation when a subclass instance should also pass — isinstance() handles that correctly, type() does not.
  • Forgetting that bool is a subclass of int, so an isinstance(x, int) check unexpectedly also accepts True/False.
  • Reaching for id() when a simple == or isinstance() check was actually what was needed.

5.11 Type Conversion

Implicit Conversion

What Is It?

Python automatically converts one type to another in certain mixed-type expressions, without you asking — most commonly, int automatically promotes to float when mixed in arithmetic.

Why Is It Used?

It avoids unnecessary manual casting for safe, lossless conversions.

>>> 1 + 2.5    # int automatically promoted to float
3.5
>>> type(1 + 2.5)
<class 'float'>

Explicit Casting Overview

How Is It Used?

For anything not automatic (like str to int), you must call the target type as a function yourself — this was already shown in 5.3 Numeric Data Types, repeated here as the general pattern:

>>> int("10") + 5
15
>>> str(15) + " items"
'15 items'

Quick Interview Answer

“Implicit conversion happens automatically in mixed-type expressions — int promotes to float in arithmetic like 1 + 2.5. Explicit conversion (casting) is anything Python won’t do on its own, like str to int — you call the target type as a function yourself: int(\"10\"), float(\"3.14\"), str(42).”

Common Mistakes

  • Expecting Python to implicitly convert a str to a number in arithmetic — it doesn’t; "10" + 5 raises a TypeError, unlike some other dynamically typed languages.
  • Forgetting that int("3.14") raises a ValueError — going from a decimal string to int requires int(float("3.14")), an explicit two-step conversion.

5.12 Memory Representation

How Python Stores Objects

Every Python object lives on the heap and carries metadata beyond its raw value — a reference count and a type pointer at minimum (the full breakdown, specific to strings, is in 9.1 Introduction to Strings).

References

What Is It?

A variable name is not a box holding a value — it’s a label pointing at an object elsewhere in memory.

Why Does It Matter?

Assigning b = a doesn’t copy a’s data; both names end up pointing at the exact same object.

flowchart LR A["a = [1, 2, 3]"] --> OBJ["[1, 2, 3]\nid: 0x7f... refcnt=2"] B["b = a"] --> OBJ

Two names referencing one shared object.

>>> a = [1, 2, 3]
>>> b = a
>>> b.append(4)
>>> a               # mutation via b is visible through a too
[1, 2, 3, 4]
>>> id(a) == id(b)
True

Garbage Collection Overview

What Is It?

CPython automatically frees an object’s memory once nothing references it anymore, primarily via reference counting (each object tracks how many references point to it; it’s freed when that count hits zero), with a supplementary cyclic garbage collector for reference cycles that counting alone can’t catch.

>>> import sys
>>> x = [1, 2, 3]
>>> sys.getrefcount(x)    # includes the temporary reference getrefcount's own argument creates
2

Quick Interview Answer

“A Python variable is a reference, not a box — b = a makes both names point at the same object rather than copying it, so mutating through one name is visible through the other. CPython frees memory primarily via reference counting: each object tracks how many references point to it and is freed at zero, with a supplementary cyclic garbage collector catching reference cycles that counting alone can’t.”

Common Mistakes

  • Assuming b = a copies a list or dict — it doesn’t; both names reference the same object, so mutating one affects the other.
  • Forgetting reference cycles exist (e.g. two objects referencing each other) — plain reference counting alone can’t free them, which is why CPython also runs a cyclic collector.

5.13 Choosing the Right Data Type

Picking the right type up front avoids both bugs and performance problems later — three questions to ask for any given piece of data.

Performance

What Is It?

Different types have very different costs for the same logical operation.

Why Does It Matter?

Checking membership (x in collection) is O(n) on a list but O(1) average on a set — for a large collection checked repeatedly, that’s the difference between a script that’s instant and one that’s noticeably slow.

# Slow for large lists: O(n) per lookup
blocked_ips = ["10.0.0.1", "10.0.0.2", ...]    # imagine thousands of entries
if user_ip in blocked_ips:
    deny_access()

# Fast: O(1) average per lookup
blocked_ips = {"10.0.0.1", "10.0.0.2", ...}    # a set instead
if user_ip in blocked_ips:
    deny_access()

Memory

A tuple generally uses less memory than an equivalent list (no room reserved for future growth), and a range() uses almost none regardless of how large the range is, since it never materializes the full sequence (see 5.4 Sequence Data Types).

Real-World Selection

SituationBest Type
Fixed collection of settings that shouldn’t changetuple
Growing/shrinking list of itemslist
Need fast “have I seen this before?” checksset
Labeled/structured data (like a JSON object)dict
Large range of numbers to iterate, not storerange
Raw file or network databytes / bytearray

Quick Interview Answer

“Choosing a data type comes down to performance, memory, and what the data actually represents. Membership testing is O(n) on a list but O(1) average on a set, so a set wins for repeated ‘have I seen this?’ checks. A tuple uses less memory than an equivalent list and signals ’this won’t change.’ range() never materializes its full sequence, so it’s nearly free memory-wise even for huge ranges.”

Common Mistakes

  • Defaulting to list for everything, including cases where a set (fast membership) or tuple (fixed, hashable record) is the actually correct choice.
  • Materializing a large range() into a list when the lazy range itself would have worked fine for iteration.

5.14 Data Types in DevOps

Python Data Types in DevOps

Real infrastructure tooling constantly maps external data (JSON APIs, config files, log lines) into these exact built-in types — here’s where each one shows up in practice.

Configuration Data

Config values loaded from a file or environment naturally become a dict — structured, labeled settings:

config = {
    "host": "localhost",
    "port": 8080,
    "debug": True,
}

JSON

json.loads() converts a JSON document directly into Python’s native types — objects become dict, arrays become list, and JSON’s true/false/null map to bool/None:

>>> import json
>>> config = json.loads('{"host":"localhost","port":8080,"debug":true,"tags":["prod","web"]}')
>>> type(config), type(config["port"]), type(config["debug"]), type(config["tags"])
(<class 'dict'>, <class 'int'>, <class 'bool'>, <class 'list'>)

API Responses

REST API responses are almost always parsed JSON — meaning nested dicts and lists, accessed the same way regardless of which API you’re calling.

response = {
    "status": "success",
    "data": {"user_id": 42, "active": True},
}
>>> response["data"]["user_id"]
42

AWS Resources

boto3 responses follow the same dict/list nesting pattern — knowing how to navigate nested dicts is directly transferable to any AWS SDK call.

instance = {
    "InstanceId": "i-0abc123",
    "State": {"Name": "running"},
    "Tags": [{"Key": "Name", "Value": "web01"}],
}
>>> instance["State"]["Name"]
'running'

Log Processing

Raw log lines are str; once parsed, the extracted fields are typically stored in a dict per line for further filtering and aggregation.

Quick Interview Answer

“In real infrastructure code, external data maps directly onto Python’s built-in types: config files and json.loads() output become dict (with nested lists), JSON’s true/false/null become bool/None, boto3 AWS responses follow the same dict/list nesting, and raw log lines start as str before being parsed into per-line dicts for filtering and aggregation.”

Common Mistakes

  • Assuming a JSON field is always present and indexing directly (response["data"]["user_id"]) instead of using .get() defensively when the API contract isn’t guaranteed.
  • Forgetting that a boto3 response nests dicts and lists arbitrarily deep — reaching for a fixed number of [...] lookups instead of walking the structure defensively.

5.15 Common Mistakes

Three type-related mistakes that catch even experienced developers off guard occasionally.

Unexpected Type Changes

Because Python is dynamically typed, reassigning a variable to a different type is silent and legal — easy to do by accident, especially after an input() call that you forgot returns str (see 4.11 Input).

Mutable Defaults

What Goes Wrong?

Using a mutable object (like []) as a function’s default argument value.

Why Is It Dangerous?

Default argument values are created ONCE, when the function is defined — not fresh on every call — so all calls that rely on the default share and accumulate into the SAME list.

>>> def add_item(item, items=[]):    # BUG: mutable default
...     items.append(item)
...     return items
...
>>> add_item("a")
['a']
>>> add_item("b")    # surprise -- 'a' is still there!
['a', 'b']

>>> def add_item_fixed(item, items=None):    # FIX: use None as sentinel
...     if items is None:
...         items = []
...     items.append(item)
...     return items
...
>>> add_item_fixed("a")
['a']
>>> add_item_fixed("b")
['b']

Comparison Pitfalls

Two common traps: floating-point values that LOOK equal but aren’t exactly, due to binary representation limits; and confusing == (equal value) with is (same object).

>>> 0.1 + 0.2 == 0.3    # NOT True -- floating point imprecision
False

>>> a = [1, 2, 3]
>>> b = [1, 2, 3]
>>> a == b, a is b      # equal VALUE, but NOT the same object
(True, False)

Quick Interview Answer

“The classic type-related gotcha is the mutable default argument: a default like items=[] is created once at function definition time, not fresh per call, so every call that relies on it shares and accumulates into the same list — fix it with items=None and create the list inside the function. The other two recurring traps are floating-point comparisons (0.1 + 0.2 != 0.3) and confusing == (value equality) with is (identity).”

Common Mistakes

  • Defining a function with a mutable default argument (def f(items=[])) instead of None as a sentinel.
  • Comparing floats with == for exact equality instead of a tolerance-based comparison (math.isclose()).
  • Using is to compare values when == was intended, or vice versa — is checks identity, == checks value equality.

5.16 Best Practices

Readability

Choose the type whose name and behavior best communicate intent — a set clearly signals “uniqueness matters here” in a way a list doesn’t, even if a list would technically also work.

Consistency

Don’t let a variable silently change type across a function’s logic (see 5.15 Common Mistakes) — if a value starts as an int, keep it an int unless converting is the explicit point of that line.

Efficient Data Structures

Default to list for an ordered, changing collection; reach for set or dict when uniqueness or fast lookup is the actual requirement — revisit 5.13 Choosing the Right Data Type’s performance guidance when a script feels slower than it should.

Quick Interview Answer

“Good data-type practice comes down to three habits: pick the type that communicates intent (a set says ‘uniqueness matters’ more clearly than a list), keep a variable’s type consistent within a function’s logic rather than letting it silently drift, and default to list for ordered collections but reach for set/dict specifically when uniqueness or fast lookup is the real requirement.”

Common Mistakes

  • Using a list everywhere out of habit, even in cases where a set or dict would communicate intent and perform better.
  • Letting a variable’s type change mid-function without a clear, intentional conversion step.

5.17 Interview Questions

Frequently Asked Questions

  • What is the difference between a list and a tuple?
  • Why is bool considered a subclass of int?
  • What’s the difference between == and is?
  • Why can’t a list be used as a dictionary key, but a tuple can?
  • What is the mutable default argument bug, and how do you fix it?
  • What is the difference between type() and isinstance()?

Scenario-Based Questions

  • You need to store a million IP addresses and check membership repeatedly and fast — which type do you choose, and why?
  • A function silently accumulates data across unrelated calls — what bug pattern should you suspect first?
  • Given a parsed JSON API response, what Python types will the JSON object, array, string, number, true/false, and null become?
  • Why might 0.1 + 0.2 == 0.3 evaluate to False, and how would you compare floats safely instead?

Quick Interview Answer

“The data-type questions that come up most are: list vs tuple (mutability, hashability), why bool is a subclass of int, == vs is (value vs identity), why only hashable/immutable types work as dict keys, and the mutable default argument bug. Scenario questions usually test whether you’d reach for a set for fast membership checks, and whether you know JSON’s true/false/null map to bool/None in Python.”

Common Mistakes

  • Answering “list and tuple are basically the same” without mentioning the mutability difference and its downstream consequences (hashability, dict-key eligibility).
  • Not being able to explain why the mutable default argument bug happens (default values are evaluated once, at function definition time), only that it’s “a gotcha.”

5.18 Hands-on Exercises

Practice Programs

  • Write a function that takes a list of mixed types and returns counts of how many are int, str, and float.
  • Given a list with duplicate IP addresses, return only the unique ones using a set.
  • Write a function demonstrating the mutable-default-argument bug, then fix it.
  • Parse a small JSON config string and print the type of each top-level value.
  • Write a script that shows id() staying the same after list.append() but changing after tuple concatenation.

Mini Projects

  • Config loader — read a dict of settings, validate each value’s type with isinstance(), and raise a clear error for any mismatch.
  • IP-address deduplicator — read a list of IPs from a (simulated) log file and report unique vs total counts using a set.
  • Type profiler — given any object, print its type, id, and whether it’s mutable or immutable (based on a lookup table of known types).

Quick Interview Answer

“These exercises reinforce the chapter’s core ideas hands-on: counting types in a mixed list exercises type()/isinstance(), deduplicating IPs exercises set, the mutable-default-argument fix exercises function default semantics, and the type profiler exercises identity (id()) plus the mutable/immutable distinction from 5.9 Mutable vs Immutable Types together.”

Common Mistakes

  • Skipping the mutable-default exercise as “just theory” — it’s one of the most common real bugs in production Python code.
  • Building the type profiler with a hardcoded if/elif chain per type instead of a lookup table, making it harder to extend.

Data Types: Chapter Practice

Mini lab

Convert the string “8080” to an integer and print its type.

Expected output: int

HintUse the chapter examples to choose the operation, then run your version before opening the solution.
Show solution
port = int("8080")
print(type(port).__name__)

Knowledge check

Does converting a value change the original string object?

Check your answerNo. int() returns an integer value; strings remain immutable.

Finish the chapter

  • Explain each operation without reading the lesson.
  • Change the sample input and predict the output before running it.
  • Revisit the chapter if the result differs from your prediction.

Back to this chapter · Choose the next chapter

6.1 Introduction to Variables

Python Variables and Memory Management
flowchart TD V["Variable"] V --> R["References\nlabel, not a box"] V --> M["Stack vs Heap\nwhere things live"] V --> C["Copying\nshallow vs deep"] V --> G["Garbage Collection\nrefcounting + cycles"] V --> S["Scope\nLEGB"] V --> F["Function Arguments\nmutable vs immutable"]

The topics this chapter covers, once the basic syntax of declaring and naming variables is already familiar.

What Is a Variable, Really?

What Is It?

Syntactically, a variable is just a name bound with = — that part is covered in 4.8 Variables. Underneath that syntax, a variable is a name bound to an object living in memory, not a container holding a value directly.

Why Does It Matter?

Whether a variable “contains” a value or merely points at one changes how assignment, function calls, and copying behave — this is the single idea the rest of this chapter builds on.

>>> server_name = "web01"
>>> print(server_name)
web01

How Is It Used?

Assign with =, then refer to the name anywhere afterward in the same scope — the mechanics of where it’s visible are covered in 6.8 Variable Scope.

Already Covered Elsewhere

This chapter assumes the basics are familiar and doesn’t re-teach them. If any of the following are new, read these first:

Real-World Examples

  • current_user = "alice" — tracking who’s logged in
  • retry_count = 0 — tracking state across a loop
  • response = requests.get(url) — holding a result for later use

Variables in DevOps Scripts

DevOps scripts lean on variables constantly to avoid hardcoding values that change per environment (expanded on in 6.11 Variables in DevOps):

region = "us-east-1"
instance_type = "t3.medium"
max_retries = 3

print(f"Launching {instance_type} in {region}")

Quick Interview Answer

“Syntactically a variable is a name bound with =, but underneath, it’s a label pointing at an object living in memory — not a box holding a value. That distinction is why assignment copies a reference rather than data, why mutable and immutable arguments behave differently in functions, and why copying a list needs copy/deepcopy instead of plain assignment.”

Common Mistakes

6.2 Objects and Variable References

Objects, Briefly

Every value a variable can refer to — numbers, strings, functions, even classes — is an object with its own identity, type, and value, covered in full in 5.2 Python Object Model. An object lives in memory independently of any name pointing at it, and is created, used, and eventually destroyed based on how many references point to it (see 6.6 Garbage Collection).

A Variable Is a Label, Not a Box

What Is It?

A variable does not directly “contain” a value the way a box would — it holds a reference (a pointer) to an object living elsewhere in memory.

Why Does It Matter?

Assignment copies the reference, never the underlying data. Two names can point at the exact same object.

flowchart LR A["a"] --> OBJ["[1, 2, 3]\nid: 0x7f... refcnt = 2"] B["b"] --> OBJ

a = [1, 2, 3] then b = a — both names label the same object; b.append(4) mutates it, visible through a too.

>>> a = [1, 2, 3]
>>> b = a                # b now references the SAME object as a
>>> id(a) == id(b)
True

Shared References

How Is It Used?

Because b = a shares the object rather than copying it, mutating through one name is visible through every other name referencing it:

>>> a = [1, 2, 3]
>>> b = a
>>> b.append(4)
>>> a                     # mutation via b is visible through a too
[1, 2, 3, 4]

This is exactly why the distinction between mutable and immutable objects matters so much — see 5.9 Mutable vs Immutable Types for the full comparison. It’s usually a bug when it’s unexpected, and it’s the root cause covered in 6.12 Common Mistakes.

Identity vs Value

Two names can hold equal values while pointing at two different objects — this is exactly why == and is answer different questions (see 5.10 Type Checking for id()/type()/isinstance() in depth):

>>> a = [1, 2, 3]
>>> b = [1, 2, 3]
>>> a == b            # equal VALUE
True
>>> a is b            # NOT the same object
False

Quick Interview Answer

“A Python variable is a label pointing at an object, not a container holding it — b = a makes b point at the exact same object as a, without copying any data. If that object is mutable, mutating it through b is visible through a too, since both names reference one shared object. == compares value; is compares identity — two objects can be equal without being the same object.”

Common Mistakes

  • Assuming b = a creates an independent copy of a list or dict — it creates a second label on the same object.
  • Using is to compare values when == is what’s actually intended — is only answers “are these the exact same object,” not “do these hold equal data.”
  • Being surprised that a function call that appends to a list argument changes the caller’s list — covered fully in 6.10 Variables in Functions.

6.3 Memory Management: Stack vs Heap

Stack vs Heap

What Is It?

The stack holds function call frames — each frame stores its local variable names and the references they point to. The heap is where the actual objects (their real data) live.

Why Does It Matter?

Understanding this split explains why passing a variable into a function passes a reference, not a full copy of the data — the deep dive on that is 6.10 Variables in Functions.

flowchart LR subgraph Stack["Stack — function call frames"] M["main()\nx -->"] G["greet()\nname -->"] end subgraph Heap["Heap — actual object data"] O1["42 (int object)"] O2["\"Alice\" (str object)"] end M --> O1 G --> O2

The stack holds references (pointers); all actual objects, mutable or immutable, live on the heap.

Memory Allocation

How Is It Used?

When you create an object — a literal, a constructor call, and so on — CPython allocates space for it on the heap and returns a reference to that location, which gets stored in whatever variable or frame is holding it.

>>> x = 42          # 42 is allocated on the heap; x (in the current frame) references it

Memory Deallocation

Once an object’s reference count drops to zero — no variable, container, or frame references it anymore — CPython frees its heap memory automatically. This is why Python has no manual free() call like C; the mechanics are covered fully in 6.6 Garbage Collection.

Object Lifetime

An object exists from creation until its last reference disappears — via reassignment, del, or the referencing scope ending (see 6.9 Lifetime of Variables) — at which point it becomes eligible for garbage collection.

Quick Interview Answer

“CPython splits memory into the stack, which holds function call frames and the references their local variables point to, and the heap, where the actual object data lives. Every object — whether an int, a list, or a custom class instance — is allocated on the heap; only the reference to it sits on the stack frame. This is why passing a variable into a function passes a reference to the same heap object, not a copy of it.”

Common Mistakes

  • Assuming Python has stack-allocated value types the way C or Go do — in CPython, every object lives on the heap; the stack only ever holds references.
  • Confusing “the stack” (call frames) with “the call stack” shown in a traceback — related, but a traceback shows the sequence of frames, not memory layout details.

6.4 Assignment Operations

Reference Assignment

b = a makes b point at the exact same object a does — no data is copied, as covered in 6.2 Objects and Variable References.

There’s No Separate “Value Assignment”

What Is It?

Python doesn’t truly have a separate “value assignment” the way some languages do — all assignment is reference assignment.

Why Does It Matter?

For an immutable value like an int, this distinction rarely matters in practice, because the value itself can’t be mutated — any operation that looks like a change actually rebinds the name to a brand-new object.

>>> c = 5
>>> d = c          # d references the same int object as c
>>> d += 1         # this REBINDS d to a new object -- doesn't mutate the int
>>> c, d
(5, 6)

Effects on Mutable Objects

How Is It Used?

The consequence of reference assignment becomes visible specifically when the shared object is mutable — append() mutates the object in place, so every name referencing it sees the change:

>>> a = [1, 2]
>>> b = a
>>> b.append(3)    # mutates the shared object in place
>>> a, b
([1, 2, 3], [1, 2, 3])

Rebind vs Mutate

OperationWhat Actually HappensVisible Through Other References?
d += 1 (int)Rebinds d to a new int objectNo — other names still point at the old value
b.append(3) (list)Mutates the existing object in placeYes — every reference to it sees the change
t = t + (3,) (tuple)Creates a brand-new tuple, rebinds tNo — see 5.9 Mutable vs Immutable Types

Quick Interview Answer

“Python has no separate ‘value assignment’ mechanism — every assignment binds a name to an object reference. d = c; d += 1 looks like it mutates c’s value, but for an immutable int it actually rebinds d to a brand-new object, leaving c untouched. The same b = a; b.append(3) pattern on a mutable list mutates the one shared object in place, so the change is visible through both a and b. The difference isn’t the assignment — it’s whether the object being pointed at is mutable.”

Common Mistakes

  • Assuming += always mutates in place — for immutable types (int, str, tuple) it rebinds; for mutable types (list, via __iadd__) it does mutate in place, which is a subtle and commonly missed distinction.
  • Forgetting that reassigning a variable never affects other names that shared its old object — only mutation does.

6.5 Copying Objects

When shared references aren’t what you want (see 6.2 Objects and Variable References), Python offers three distinct ways to copy a mutable object, each with a different depth.

flowchart TD subgraph Shallow["Shallow Copy — copy.copy(original)"] O1["original"] --> L1["[1, 2, ...] (new outer list)"] S1["shallow"] --> L1 L1 --> N1["[3, 4] (SHARED nested list)"] end subgraph Deep["Deep Copy — copy.deepcopy(original)"] O2["original"] --> L2["[1, 2, ...]"] D2["deep"] --> L3["[1, 2, ...] (own copy)"] L2 --> N2["[3, 4] (original)"] L3 --> N3["[3, 4] (own copy)"] end

Shallow copy duplicates the outer container but shares nested objects; deep copy duplicates everything, recursively.

Assignment Copy

b = a is not a copy at all — both names reference the identical object. Included here only to contrast with the two real copy types below.

Shallow Copy

What Is It?

Creates a new outer object, but the elements inside it are still shared references to the same nested objects as the original.

Why Is It Used?

Fast, and sufficient when the container holds only immutable elements, or nested mutation isn’t a concern.

Deep Copy

What Is It?

Recursively copies the object and everything it contains, so the result is completely independent of the original at every level.

Why Is It Used?

Needed whenever the original contains nested mutable objects (like a list of lists) and full independence is required.

The copy Module

The standard library’s copy module provides both copy() and deepcopy() — here’s all three approaches compared directly:

>>> import copy
>>> original = [1, 2, [3, 4]]
>>> assign_copy = original
>>> shallow = copy.copy(original)
>>> deep = copy.deepcopy(original)

>>> original[2].append(99)      # mutate the nested list

>>> original
[1, 2, [3, 4, 99]]
>>> assign_copy                 # same object -- sees the change
[1, 2, [3, 4, 99]]
>>> shallow                     # outer copied, but nested list is SHARED -- sees it
[1, 2, [3, 4, 99]]
>>> deep                        # fully independent -- does NOT see the change
[1, 2, [3, 4]]

Quick Interview Answer

“b = a isn’t a copy at all — both names share one object. copy.copy() creates a shallow copy: a new outer container, but its nested elements are still shared with the original, so mutating a nested list through either one is visible through both. copy.deepcopy() recursively copies everything, producing a result that’s fully independent at every level. Use shallow copy when the container only holds immutable elements; use deep copy whenever it holds nested mutable objects and true independence is required.”

Common Mistakes

  • Assuming copy.copy() produces a fully independent object — it only copies the outer container; nested mutable objects are still shared.
  • Reaching for copy.deepcopy() by default even when not needed — it’s slower and unnecessary for flat structures or containers of immutable values.
  • Using list(original) or original[:] expecting a deep copy — both are shallow copies, identical in depth to copy.copy().

6.6 Garbage Collection

Reference Counting

What Is It?

CPython’s primary memory-management mechanism, introduced in 5.12 Memory Representation — every object tracks how many references point to it, incrementing on each new reference and decrementing when one goes away.

Why Is It Used?

As soon as this count hits zero, the object’s memory is freed immediately, without waiting for a periodic sweep — unlike garbage collectors in some other languages.

>>> import sys
>>> x = [1, 2, 3]
>>> sys.getrefcount(x)   # includes the temporary reference getrefcount's own call creates
2

The Cyclic Garbage Collector

What Is It?

A supplementary mechanism that specifically detects and cleans up reference cycles — e.g. two objects that reference each other — which reference counting alone can never resolve, since each one’s count never naturally reaches zero.

>>> import gc
>>> gc.isenabled()
True
>>> gc.collect()      # force a collection pass; returns count of objects collected
0

del and Memory Cleanup

del removes a name’s reference to an object (see also 6.9 Lifetime of Variables); if that was the last reference, the object becomes eligible for immediate cleanup via reference counting.

>>> y = [1, 2, 3]
>>> del y
>>> y
Traceback (most recent call last):
NameError: name 'y' is not defined

Quick Interview Answer

“CPython frees memory through two mechanisms working together. Reference counting is primary: every object tracks how many references point to it, and it’s freed the instant that count hits zero — no waiting for a sweep. That alone can’t free reference cycles, like two objects pointing at each other, since neither’s count ever naturally reaches zero — so a supplementary cyclic garbage collector, exposed through the gc module, periodically scans for and cleans up exactly that case. del just removes one name’s reference; the object is only actually freed once its refcount reaches zero.”

Common Mistakes

  • Believing Python has no garbage collector because reference counting handles most cases — the cyclic collector is a real, necessary second mechanism for reference cycles.
  • Thinking del x immediately destroys the object — it only removes that one reference; the object survives if anything else still references it.
  • Calling gc.collect() routinely in application code “just in case” — it’s rarely needed outside of debugging memory issues or specific long-running-process tuning.

6.7 Memory Optimization

Object Reuse

Where possible, CPython reuses existing immutable objects instead of allocating new ones for identical values. The two mechanisms below are the main examples.

String Interning

CPython automatically shares one object for many identical, identifier-like string literals:

>>> s1 = "hello"
>>> s2 = "hello"
>>> s1 is s2       # same interned object
True

Integer Caching

What Is It?

CPython pre-creates and reuses integer objects for the small range -5 to 256 — any int in that range, no matter how it’s produced, refers to the same cached object.

Why Is It Used?

These are extremely common values (loop counters, small flags), so caching them avoids constant reallocation.

>>> a = 100
>>> b = 100
>>> a is b         # within the cached range (-5 to 256)
True

>>> x = 1000
>>> y = 2000
>>> z = y - 1000   # computed at runtime, same value as x
>>> z is x         # outside the cached range -- NOT guaranteed to match
False

Important: Never rely on integer caching or string interning for correctness — these are CPython implementation details, not language guarantees. Always compare values with ==, and reserve is for genuine identity checks (e.g. against None).

Efficient Variable Usage

  • Reuse existing variables instead of creating many short-lived duplicates in tight loops.
  • Prefer built-in operations (join(), comprehensions) over manual accumulation where possible.
  • Delete large objects explicitly (del) once truly done with them in long-running processes.

Quick Interview Answer

“CPython applies two main object-reuse optimizations: string interning, which shares one object for identical, identifier-like string literals, and small integer caching, which pre-creates and reuses int objects for -5 to 256. Both are why is can appear to work for value comparison on small values in a REPL — but neither is a language guarantee, only a CPython implementation detail, so correctness code should always compare with == and reserve is for genuine identity checks like x is None.”

Common Mistakes

  • Using is to compare two integers or strings for equality because it “worked in testing” — it only appears to work inside the cached/interned range, and breaks silently outside it.
  • Assuming these optimizations apply the same way across all Python implementations — they’re specific to CPython’s implementation, not part of the language specification.

6.8 Variable Scope

flowchart TD B["Built-in\nlen, print, range, ... — always available"] Gs["Global\nnames at module level"] E["Enclosing\nnames in an outer (enclosing) function"] L["Local\nnames inside the current function"] L --> E --> Gs --> B

Python searches Local → Enclosing → Global → Built-in, stopping at the first match.

Local Scope

What Is It?

Names created inside a function exist only within that function call, and disappear once it returns.

Why Is It Used?

Keeps a function’s internal working variables from leaking into or colliding with the rest of the program.

def greet():
    message = "Hello"    # local to greet()
    print(message)

>>> greet()
Hello
>>> message                # not accessible outside the function
Traceback (most recent call last):
NameError: name 'message' is not defined

Global Scope

Names defined at module (top) level are accessible for reading from anywhere in that module, including inside functions — but writing to them from inside a function requires the global keyword.

The global Keyword

What Is It?

Tells Python that an assignment inside a function should modify the module-level (global) variable, instead of creating a new local one. Without it, any assignment to that name inside the function creates a separate local variable — the exact cause of the UnboundLocalError bug covered in 6.12 Common Mistakes.

count = 0

def increment():
    global count
    count += 1

>>> increment(); increment()
>>> count
2

nonlocal

The equivalent of global, but for reaching one level up into an enclosing function’s scope (not all the way to module level) — used in nested functions/closures.

def outer():
    x = "enclosing"
    def inner():
        nonlocal x
        x = "modified by inner"
    inner()
    print(x)

>>> outer()
modified by inner

The LEGB Rule

The order Python searches when resolving a name: Local → Enclosing → Global → Built-in, stopping at the first scope where the name is found. This explains exactly which variable a given name refers to in any nested-function situation.

Quick Interview Answer

“Python resolves any name using the LEGB rule: Local scope first, then any Enclosing function’s scope, then module-level Global scope, then Built-ins — stopping at the first match. Reading a global from inside a function works without any keyword, but writing to it requires global, or Python creates a new local variable instead. nonlocal is the same idea one level up, for reaching into an enclosing function’s scope from a nested closure rather than all the way to module level.”

Common Mistakes

  • Forgetting global when trying to reassign a module-level variable inside a function — the assignment silently creates a new local variable instead, and the global is left untouched.
  • Confusing global and nonlocal — global always reaches module level; nonlocal reaches exactly one enclosing function scope, not all the way up.
  • Assuming Python has block scope (like if/for) — it doesn’t; see 6.12 Common Mistakes for the “incorrect scope” bug this causes.

6.9 Lifetime of Variables

Creation

A variable’s lifetime begins the moment it’s first assigned a value.

Usage

A variable remains usable for as long as its scope is active (see 6.8 Variable Scope) and it hasn’t been deleted.

Deletion

A variable’s lifetime ends when its scope exits — a function returns, for instance — or it’s explicitly removed with del.

The del Statement

What Is It?

Explicitly removes a name’s binding, immediately. If that was the object’s last reference, it becomes eligible for garbage collection right away (see 6.6 Garbage Collection).

>>> y = [1, 2, 3]
>>> del y
>>> y
Traceback (most recent call last):
NameError: name 'y' is not defined

How Is It Used?

del removes the name, not necessarily the object — if another variable still references the same object, that object stays alive:

>>> a = [1, 2, 3]
>>> b = a
>>> del a
>>> b            # the object is still alive -- b still references it
[1, 2, 3]

Quick Interview Answer

“A variable’s lifetime runs from its first assignment until its scope exits or it’s explicitly removed with del. del only removes that one name’s binding, immediately — it doesn’t necessarily destroy the underlying object, which stays alive as long as any other reference to it exists. The object itself is only freed once its reference count reaches zero, which is a separate concern covered by garbage collection.”

Common Mistakes

  • Assuming del x destroys the object x referenced — it only removes that one name’s binding; the object survives if anything else still references it.
  • Relying on a function-local variable’s value surviving after the function returns — its lifetime ends with the function call unless the value was explicitly returned.

6.10 Variables in Functions

Passing Variables

What Is It?

Python passes arguments by “assignment” — the parameter name inside the function becomes a new reference to the same object the caller passed in. It’s neither a copy of the data (unlike C’s pass-by-value) nor a true reference-to-the-variable (unlike C++’s pass-by-reference); this model is sometimes called “pass by object reference.”

Why Does It Matter?

Whether changes inside the function are visible to the caller depends entirely on whether the argument’s type is mutable, not on any special syntax — building directly on 6.4 Assignment Operations.

Mutable vs Immutable Arguments

def modify_list(lst):
    lst.append(4)     # mutates the SAME object the caller has

def modify_int(n):
    n += 1             # rebinds the LOCAL name n -- caller's variable is untouched
    return n

>>> l = [1, 2, 3]
>>> modify_list(l)
>>> l                  # caller's list WAS mutated
[1, 2, 3, 4]

>>> num = 5
>>> result = modify_int(num)
>>> num, result         # caller's int was NOT changed; a new value was returned instead
(5, 6)

Return Values

Since reassignment inside a function never affects the caller’s variable for immutable types, the standard pattern is to explicitly return the new value and have the caller reassign it if needed — exactly as result = modify_int(num) does above.

Quick Interview Answer

“Python passes arguments by object reference — the parameter becomes a new name pointing at the same object the caller passed in. Whether the function’s changes are visible to the caller depends purely on mutability: mutating a mutable argument in place (lst.append(4)) is visible to the caller, since it’s the same object; reassigning an immutable argument (n += 1) just rebinds the local parameter name to a new object, leaving the caller’s variable untouched. That’s why functions dealing with immutable values follow the return-and-reassign pattern instead.”

Common Mistakes

  • Describing Python as “pass by value” or “pass by reference” — it’s neither; it’s pass-by-object-reference, and the visible behavior depends on the argument’s mutability.
  • Expecting a function like modify_int above to change the caller’s variable just because it looks similar to modify_list — only mutation is visible to the caller, not reassignment.
  • Relying on mutating a passed-in list as a way to “return” a result instead of using an explicit return — works, but is far less readable and easy to misuse.

6.11 Variables in DevOps

DevOps work is full of environment-specific values — variables, in every sense (Python, OS environment, and Terraform), are how they’re managed instead of hardcoded.

Configuration Variables

Settings that differ per deployment (host, port, feature flags), typically loaded from a file or environment rather than hardcoded:

config = {
    "host": "localhost",
    "port": 8080,
    "debug": False,
}

Environment Variables

What Are They?

Values set in the OS environment, outside the Python script itself — the standard way to inject deployment-specific config (like which environment: dev/staging/prod) without changing code.

>>> import os
>>> os.environ.get("HOME")
'/root'
>>> os.environ.get("NOT_SET", "default_value")
'default_value'

Secrets

What Is It?

Sensitive values (API keys, passwords, tokens) that must never be hardcoded into source code.

Why Does It Matter?

Hardcoded secrets end up in version control history permanently — even if removed in a later commit, they remain recoverable from the git history.

How Is It Used?

Load from environment variables or a dedicated secrets manager (AWS Secrets Manager, HashiCorp Vault) at runtime instead:

import os

api_key = os.environ.get("API_KEY")   # never: api_key = "sk-abc123..."
if not api_key:
    raise RuntimeError("API_KEY environment variable not set")

AWS Credentials

boto3 automatically reads AWS credentials from environment variables (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY) or the shared ~/.aws/credentials file — a script’s own variables typically never touch the raw credentials directly:

import boto3

# boto3 picks up credentials from the environment/config automatically --
# no credentials appear as literal variables in the script itself
s3 = boto3.client("s3")

Terraform Variables

Terraform has its own variable system (variable blocks in .tf files), conceptually similar to Python variables — named, typed placeholders for values that differ per environment, such as instance_type or region.

Quick Interview Answer

“DevOps scripts lean on three layers of variables: Python variables for in-script logic, OS environment variables (os.environ) for deployment-specific config injected without changing code, and infrastructure-as-code variables like Terraform’s variable blocks for the same idea at the infrastructure layer. Secrets specifically should never be hardcoded as Python literals — they belong in environment variables or a dedicated secrets manager, since a hardcoded secret remains recoverable from git history even after being removed in a later commit.”

Common Mistakes

  • Hardcoding an API key or password directly as a string literal — it persists in version control history indefinitely, even after later “removing” it.
  • Not providing a default with os.environ.get(name, default) and instead using os.environ[name], which raises an unhandled KeyError if the variable isn’t set.
  • Assuming boto3 needs credentials passed explicitly as variables in code — it resolves them automatically from the environment or ~/.aws/credentials unless overridden.

6.12 Common Mistakes

UnboundLocalError

What Goes Wrong?

Referencing a variable inside a function before assigning it, when that same name is also assigned somewhere later in that same function — Python decides the name is local to the whole function (because it’s assigned there), so the earlier read fails.

count = 5

def broken():
    print(count)     # Python sees the later assignment below and treats count as LOCAL here
    count = 1

>>> broken()
Traceback (most recent call last):
UnboundLocalError: cannot access local variable 'count' where it is not associated with a value

The fix is the global keyword covered in 6.8 Variable Scope.

Shared Mutable Objects

Forgetting that b = a shares the same object (see 6.2 Objects and Variable References), then being surprised when mutating b also changes a. The mutable-default-argument variant of this bug — a default argument evaluated once, at function-definition time, and shared across every call — is one of the most common real-world instances of this pattern.

Variable Shadowing

What Is It?

A local variable (or function parameter) using the same name as a global variable or built-in function, silently hiding it within that scope.

list = [1, 2, 3]      # shadows the built-in list() constructor!
>>> list("abc")        # this now fails -- 'list' is a variable, not the builtin
Traceback (most recent call last):
TypeError: 'list' object is not callable

Incorrect Scope

What Goes Wrong?

Assuming a variable assigned inside an if or for block is scoped to just that block — Python has no block scope, only function scope, so it actually leaks into the surrounding function or module.

if True:
    x = 10             # NOT block-scoped -- this leaks into the enclosing scope

>>> print(x)             # accessible here, unlike in C/Java
10

Quick Interview Answer

“The recurring variable-related bugs are: UnboundLocalError, from reading a name before an assignment further down in the same function makes Python treat it as local; shared mutable objects, from forgetting b = a shares one object rather than copying it; variable shadowing, from naming a local the same as a builtin like list; and assuming block scope exists for if/for, when Python only has function scope, so a variable assigned inside a loop or conditional leaks into the surrounding function.”

Common Mistakes

  • Fixing an UnboundLocalError by renaming the variable instead of understanding why it happened — the real fix is usually global, or restructuring so the read doesn’t precede the write.
  • Naming a variable list, dict, str, type, or id — all shadow built-ins and cause confusing failures much later in the same scope.
  • Relying on a loop variable’s final value after the loop ends — it works because there’s no block scope, but it’s fragile if the loop body is later restructured.

6.13 Best Practices

Meaningful Names

Choose names that describe what a variable holds (user_count, not uc) — this is the single highest-leverage readability habit, covered in 4.7 Identifiers.

Avoid Global Variables

Prefer passing values as function arguments and returning results over reading/writing global state — globals make code harder to test and reason about, since any function can silently change them (see the UnboundLocalError and global-keyword pitfalls in 6.12 Common Mistakes).

Use Constants

Replace magic numbers and strings scattered through code with named UPPER_SNAKE_CASE constants defined once — see 4.9 Constants — easier to update and self-documenting.

Memory-Efficient Coding

  • Prefer generators/iterators over building large intermediate lists when possible.
  • Delete large objects explicitly (del) once truly done with them in long-running scripts — see 6.6 Garbage Collection.
  • Reach for a shallow copy when it’s sufficient, and reserve copy.deepcopy() for nested mutable structures that genuinely need full independence — see 6.5 Copying Objects.
  • Never rely on integer caching or string interning for correctness — see 6.7 Memory Optimization.

Quick Interview Answer

“The practical habits that matter most for variables and memory: use descriptive names over abbreviations, avoid mutable global state in favor of passing arguments and returning results, replace magic values with named constants, and — for long-running or memory-sensitive scripts — prefer generators over building large lists, copy only as deep as actually needed, and explicitly del large objects once done with them.”

Common Mistakes

  • Treating these as optional style preferences rather than habits that prevent real bugs — global state and shared mutable objects are two of the most common sources of production incidents.
  • Optimizing memory prematurely in small scripts where readability matters far more than shaving a few allocations.

6.14 Interview Questions

Frequently Asked Questions

  • What is the difference between a variable and an object in Python?
  • Explain the difference between is and == with an example.
  • What is the difference between a shallow copy and a deep copy?
  • What causes an UnboundLocalError, and how do you fix it?
  • Explain the LEGB rule.
  • Why does Python need both reference counting and a garbage collector?

Scenario-Based Questions

  • A function receives a list, appends to it, and the caller sees the change — but a similar function with an int argument doesn’t affect the caller. Why?
  • You copy a list of lists with copy.copy() and mutate a nested list — the “copy” changes too. What copy type do you need instead, and why?
  • A script sets an API key directly in the source code. What’s wrong with this, and what’s the fix?

Quick Interview Answer

“The core questions in this area test whether you understand that a variable is a reference, not a container: is vs == (identity vs value), shallow vs deep copy (whether nested objects are shared), UnboundLocalError (a name assigned anywhere in a function is treated as local throughout it), the LEGB scope-resolution order, and why reference counting alone can’t free reference cycles, which is what the supplementary cyclic garbage collector exists for. The scenario questions usually come down to one root cause: mutable objects are shared through references, immutable ones are not.”

Common Mistakes

  • Answering “Python passes by reference” or “by value” for function arguments instead of the more precise “pass by object reference” — see 6.10 Variables in Functions.
  • Explaining shallow vs deep copy only in the abstract, without being able to walk through what happens to a nested list under each — see 6.5 Copying Objects.
  • Not being able to explain why UnboundLocalError happens (Python decides a name is local for the whole function if it’s assigned anywhere in it), only that it’s “a scoping bug.”

6.15 Hands-on Exercises

Practice Programs

  • Write a function demonstrating UnboundLocalError, then fix it using the global keyword.
  • Given a nested list, show the difference between copy.copy() and copy.deepcopy() by mutating a nested element.
  • Write a closure using nonlocal to implement a simple counter.

Mini Projects

  • Configuration loader — build a small config loader that reads settings from environment variables with sensible defaults, using 6.11 Variables in DevOps as a starting point.
  • Reference tracer — build a tool that takes any object and prints its id(), type(), and sys.getrefcount() before and after a mutation, to make the concepts from 6.2 Objects and Variable References and 6.6 Garbage Collection concrete.

Quick Interview Answer

“These exercises reinforce the chapter’s core ideas hands-on: the UnboundLocalError-then-fix exercise exercises scope and the global keyword; the shallow-vs-deep-copy exercise exercises reference sharing on nested mutable objects; the nonlocal counter exercises closures and enclosing scope; and the reference tracer ties identity, type, and reference counting together into one small diagnostic tool.”

Common Mistakes

  • Skipping the shallow-vs-deep-copy exercise as “just theory” — it’s one of the most common real bugs when working with nested config dicts or lists of records.
  • Building the reference tracer without accounting for the temporary reference sys.getrefcount()’s own call creates, and misreading the count as a result.

Variables and Memory: Chapter Practice

Mini lab

Copy a list before appending a server name. Show that the original is unchanged.

Expected output: [‘web-01’] and then [‘web-01’, ‘db-01’]

HintUse the chapter examples to choose the operation, then run your version before opening the solution.
Show solution
original = ["web-01"]
copy = original.copy()
copy.append("db-01")
print(original)
print(copy)

Knowledge check

Does a shallow copy copy nested lists too?

Check your answerNo. Nested mutable values are still shared.

Finish the chapter

  • Explain each operation without reading the lesson.
  • Change the sample input and predict the output before running it.
  • Revisit the chapter if the result differs from your prediction.

Back to this chapter · Choose the next chapter

7.1 Introduction to Operators

Python Operators
flowchart TD O["Operators"] O --> AR["Arithmetic\n+ - * / // % **"] O --> AS["Assignment\n= += -= ..."] O --> CO["Comparison\n== != < > <= >="] O --> LO["Logical\nand or not"] O --> BW["Bitwise\n& | ^ ~ << >>"] O --> ME["Membership\nin / not in"] O --> ID["Identity\nis / is not"]

The seven operator categories this chapter covers.

What Are Operators?

What Is It?

An operator is a symbol (+, ==, and, …) that performs an operation on one or more values, producing a result.

Why Does It Matter?

Operators are the basic building blocks of every calculation, comparison, and condition a program makes.

How Is It Used?

Combined with operands to form expressions, evaluated according to precedence rules (see 7.9 Operator Precedence and Expression Evaluation).

>>> 5 + 3
8

Operands

The values an operator acts on. In 5 + 3, both 5 and 3 are operands of the + operator. Operands can be literals, variables, or entire sub-expressions.

Expressions

What Is It?

Any combination of operators and operands that evaluates to a value. Every expression produces exactly one result, which can itself become an operand in a larger expression.

>>> x = 5
>>> (x + 3) * 2      # the whole thing is one expression
16

Why Operators Matter

Nearly every meaningful line of code — a calculation, a condition, a validity check — relies on operators. Understanding their exact behavior, especially precedence and short-circuiting, prevents an entire category of subtle logic bugs.

Real-World DevOps Use Cases

  • Comparing CPU/disk usage against alert thresholds
  • Checking HTTP status codes fall in a valid range
  • Combining multiple health-check conditions with and/or
  • Bit flags for permissions or feature toggles

Expanded fully in 7.16 DevOps Use Cases.

Quick Interview Answer

“An operator is a symbol that performs an operation on one or more operands, producing a result — combined together, operators and operands form expressions. Python groups its operators into categories: arithmetic, assignment, comparison, logical, bitwise, membership, and identity. Getting precedence and short-circuit evaluation right is what separates code that merely looks correct from code that actually is.”

Common Mistakes

  • Treating operator behavior as “obvious” and skipping precedence entirely — 2 + 3 * 4 silently evaluates to 14, not 20, which is exactly the kind of bug that passes a quick visual check.
  • Assuming every operator means the same thing regardless of type — + is arithmetic addition for numbers but concatenation for sequences, covered in 7.13 Operators with Different Data Types.

7.2 Arithmetic Operators

The standard mathematical operators, all usable on int and float (and some on other types, see 7.13 Operators with Different Data Types).

Addition (+)

>>> 5 + 3
8

Subtraction (-)

>>> 5 - 3
2

Multiplication (*)

>>> 5 * 3
15

Division (/)

What Is It?

True division — always returns a float, even if the result is a whole number.

>>> 5 / 3
1.6666666666666667

Floor Division (//)

What Is It?

Divides and rounds down to the nearest whole number.

Why Is It Used?

When a whole-number result is needed — e.g. “how many full batches of 3 fit into 5 items?”

>>> 5 // 3
1

Modulus (%)

What Is It?

Returns the remainder of division.

Why Is It Used?

Extremely common for “every Nth item” logic, and checking even/odd.

>>> 5 % 3
2

Exponent (**)

>>> 5 ** 3
125

Practical Example

>>> total_items = 17
>>> batch_size = 5
>>> full_batches = total_items // batch_size
>>> leftover = total_items % batch_size
>>> full_batches, leftover
(3, 2)

Quick Interview Answer

“Python’s arithmetic operators are + - * / // % **. The one that trips people up most is / vs //: / is true division and always returns a float, while // is floor division and rounds down to a whole number. % returns the remainder, which combined with // is the standard pattern for splitting a total into full batches plus a leftover.”

Common Mistakes

  • Expecting / to return an int when both operands are int — it always returns a float in Python 3, unlike Python 2’s /, covered further in 7.14 Common Mistakes.
  • Confusing // with rounding to the nearest whole number — it always rounds down (toward negative infinity), not to the nearest value.
  • Forgetting ** is right-associative — 2 ** 3 ** 2 is 2 ** (3 ** 2) = 512, not (2 ** 3) ** 2, detailed in 7.9 Operator Precedence and Expression Evaluation.

7.3 Assignment Operators

=

Plain assignment binds a name to a value — already covered in 4.8 Variables and 6.4 Assignment Operations.

>>> x = 5
>>> x
5

Augmented Assignment

What Is It?

Shorthand combining an arithmetic/bitwise operation with assignment in one step — x += 1 means x = x + 1.

Why Is It Used?

Less repetition, and the intent (“update this variable”) is clearer at a glance.

OperatorEquivalent ToExample (starting x=5)Result
+=x = x + nx += 38
-=x = x - nx -= 2 (from 8)6
*=x = x * nx *= 2 (from 6)12
/=x = x / nx /= 4 (from 12)3.0
//=x = x // nx //= 3 (10 ->)3
%=x = x % nx %= 2 (from 3)1
**=x = x ** nx **= 4 (2 ->)16
&=x = x & nx &= 3 (5 ->)1
|=x = x | nx |= 2 (5 ->)7
^=x = x ^ nx ^= 1 (5 ->)4
<<=x = x << nx <<= 3 (1 ->)8
>>=x = x >> nx >>= 2 (16 ->)4

Quick Interview Answer

“Augmented assignment operators like +=, -=, *= combine an operation with assignment in one step — x += 3 is shorthand for x = x + 3. They exist for both arithmetic (+= -= *= /= //= %= **=) and bitwise operators (&= |= ^= <<= >>=). For an immutable value like an int, x += 1 rebinds x to a new object; for a mutable value like a list, x += [4] mutates the existing object in place via __iadd__ — the same rebind-vs-mutate distinction covered in 6.4 Assignment Operations.”

Common Mistakes

  • Assuming x += 1 always mutates in place — for immutable types it rebinds to a new object, which only matters when another variable shares the old reference.
  • Chaining augmented assignment across unrelated statements and losing track of the running value — prefer clear, separate steps when the sequence isn’t obvious at a glance.

7.4 Comparison Operators

What Are They?

Operators that compare two values and always return a bool.

Why Are They Used?

The foundation of every conditional (if, while) in a program.

OperatorMeaningExampleResult
==Equal to5 == 5True
!=Not equal to5 != 3True
<Less than5 < 3False
>Greater than5 > 3True
<=Less than or equal5 <= 5True
>=Greater than or equal5 >= 6False

Comparison Chaining

What Is It?

Writing a < b < c evaluates as (a < b) and (b < c) — both comparisons must hold, and b is only evaluated once. Covered in more depth in 7.12 Chained Comparisons.

>>> 1 < 2 < 3
True

Version Comparison

What Is It?

Comparing tuples element-by-element — exactly how version numbers like (1, 5, 0) are typically compared, since Python compares tuples lexicographically (the first differing element decides the result).

>>> v1 = (1, 5, 0)
>>> v2 = (1, 4, 9)
>>> v1 > v2      # compares 1==1, then 5>4 decides it
True

Quick Interview Answer

“Comparison operators (== != < > <= >=) always return a bool and are the foundation of every if/while condition. Python allows chaining them directly — a < b < c — which evaluates as (a < b) and (b < c), with b evaluated only once. A practical use of comparison is comparing tuples lexicographically, which is exactly how version tuples like (1, 5, 0) are typically compared.”

Common Mistakes

  • Using == on floats expecting exact equality — floating-point rounding means 0.1 + 0.2 == 0.3 is False; compare with a tolerance instead (abs(a - b) < 1e-9).
  • Using is where == was intended — is compares identity, not value, covered fully in 7.8 Identity Operators.

7.5 Logical Operators

flowchart TD subgraph AND["and — True only if BOTH are truthy"] direction LR A1["True and True → True"] A2["True and False → False"] A3["False and True → False"] A4["False and False → False"] end subgraph OR["or — True if AT LEAST ONE is truthy"] direction LR O1["True or True → True"] O2["True or False → True"] O3["False or True → True"] O4["False or False → False"] end subgraph NOT["not — inverts"] direction LR N1["not True → False"] N2["not False → True"] end

Truth tables for and, or, and not.

and

Returns True only if both operands are truthy.

>>> True and False
False

or

Returns True if at least one operand is truthy.

>>> True or False
True

not

Inverts a boolean value.

>>> not True
False

Short-Circuit Evaluation

What Is It?

and/or stop evaluating as soon as the overall result is already determined by the first operand — the second operand is never even executed in that case.

Why Does It Matter?

Commonly exploited to avoid errors (checking x is not None and x.value safely) or to avoid unnecessary work.

flowchart LR subgraph AndCase["A and B — A = False"] A1["A = False"] -->|"short-circuits"| R1["Result: False\n(B never evaluated)"] end subgraph OrCase["A or B — A = True"] A2["A = True"] -->|"short-circuits"| R2["Result: True\n(B never evaluated)"] end

and stops at the first False; or stops at the first True.

def side_effect():
    print("called")
    return True

>>> False and side_effect()    # side_effect() is NEVER called
False
>>> True or side_effect()      # side_effect() is NEVER called
True

Quick Interview Answer

“and returns True only if both operands are truthy; or returns True if at least one is; not inverts a boolean. Both and and or short-circuit — they stop evaluating as soon as the result is already determined, so the second operand may never actually execute. This is exploited deliberately to guard against errors, like x is not None and x.value, where x.value is only evaluated once x is not None has already confirmed it’s safe.”

Common Mistakes

  • Assuming both operands of and/or are always evaluated — short-circuiting means a function call used as the second operand may silently never run.
  • Writing if x == True: or if x == False: instead of if x: or if not x: — un-Pythonic and breaks for truthy/falsy values that aren’t literally True/False (see 7.10 Boolean Evaluation).

7.6 Bitwise Operators

What Are They?

Operators that work on the individual bits of an integer’s binary representation, rather than its decimal value.

Why Are They Used?

Permission flags, low-level protocol parsing, and performance-sensitive bit manipulation.

Binary Representation

Every int has an underlying binary form — bin() displays it, and format() lets you control the padding:

>>> bin(5), bin(3)
('0b101', '0b11')
>>> format(5, '04b')     # zero-padded to 4 bits
'0101'

&, |, and ^ applied bit-by-bit to 5 (0101) and 3 (0011):

Bit positiona=5 (0101)b=3 (0011)a & ba | ba ^ b
bit 300000
bit 210011
bit 101011
bit 011110
Result0001 = 10111 = 70110 = 6

& (AND)

Each result bit is 1 only if both input bits are 1.

>>> 5 & 3
1

| (OR)

Each result bit is 1 if either input bit is 1.

>>> 5 | 3
7

^ (XOR)

Each result bit is 1 if the input bits differ.

>>> 5 ^ 3
6

~ (NOT)

Inverts every bit. In Python’s signed representation, this is equivalent to -(x + 1).

>>> ~5
-6

<< (Left Shift)

Shifts bits left, equivalent to multiplying by 2 per shifted position.

>>> 5 << 1
10

>> (Right Shift)

Shifts bits right, equivalent to floor-dividing by 2 per shifted position.

>>> 5 >> 1
2

Bit Manipulation Example

A classic real-world use: representing a set of independent permission flags in a single integer, using one bit per flag.

>>> READ, WRITE, EXECUTE = 4, 2, 1     # 100, 010, 001
>>> perms = READ | WRITE               # combine flags
>>> bin(perms)
'0b110'
>>> bool(perms & READ)                 # check if READ flag is set
True
>>> bool(perms & EXECUTE)              # check if EXECUTE flag is set
False

Quick Interview Answer

“Bitwise operators work on an integer’s binary representation rather than its decimal value: & (AND) and | (OR) work bit-by-bit like their logical counterparts, ^ (XOR) is true only where bits differ, ~ inverts every bit (equivalent to -(x+1) in Python’s signed representation), and <</>> shift bits left/right, equivalent to multiplying or floor-dividing by a power of two. The classic real-world use is packing independent boolean flags — like permissions — into a single integer, one bit per flag, then testing a flag with perms & FLAG.”

Common Mistakes

  • Confusing bitwise &/| with logical and/or — bitwise operators work bit-by-bit on integers (or set union/intersection, see 7.13 Operators with Different Data Types); logical operators work on truthy/falsy values as a whole.
  • Expecting ~5 to be -5 — Python’s two’s-complement-style signed representation makes it -(x + 1), so ~5 is -6.
  • Reaching for bit flags in application code where a set of named booleans or an Enum with flags would be far more readable.

7.7 Membership Operators

in

What Is It?

Tests whether a value exists within a collection (list, string, dict keys, and so on), returning a bool.

>>> 3 in [1, 2, 3]
True
>>> "a" in "cat"
True
>>> "key" in {"key": 1}    # checks dict KEYS by default
True

not in

The negation of in — tests for absence.

>>> 5 not in [1, 2, 3]
True

Searching Collections

Performance Note

in is O(n) on a list but O(1) average on a set or dict — see 5.6 Set Data Types. Convert to a set first if checking membership repeatedly against a large collection.

Validation Example

allowed_regions = {"us-east-1", "us-west-2", "eu-west-1"}
region = "us-east-1"

>>> region in allowed_regions
True

Quick Interview Answer

“in and not in test whether a value is present in (or absent from) a collection, returning a bool — checking dict membership tests keys by default. The performance characteristics differ sharply by container: in is O(n) on a list, since it has to scan element by element, but O(1) average on a set or dict, since both use hashing under the hood. Any code doing repeated membership checks against a large collection should convert it to a set first.”

Common Mistakes

  • Repeatedly checking x in some_list inside a loop against a large list instead of converting it to a set once beforehand — an easy, invisible performance cliff at scale.
  • Forgetting in on a dict checks keys, not values — use value in d.values() explicitly if that’s what’s actually needed.

7.8 Identity Operators

This section covers is/is not specifically as comparison operators. The underlying concept — a variable being a reference to an object, not the object itself — is covered in full in 6.2 Objects and Variable References.

is

What Is It?

Tests whether two names reference the exact same object in memory (same identity), not merely equal values.

>>> a = [1, 2, 3]
>>> b = a
>>> c = [1, 2, 3]
>>> a is b      # same object
True
>>> a is c      # equal value, but a DIFFERENT object
False

is not

>>> a is not c
True

Identity vs Equality

The Rule of Thumb

Use == to compare values (almost always what you want), and reserve is specifically for identity checks — most commonly x is None, since None is a singleton.

>>> a == c    # equal value
True

id()

Directly inspecting the identities being compared — see also 5.10 Type Checking:

>>> id(a), id(c)
(140475706460992, 140475706462784)    # different -- confirms a is c is False

Quick Interview Answer

“is tests identity — whether two names point at the exact same object — while == tests value equality. Two lists can hold identical values (== is True) while being two entirely separate objects (is is False). The standard rule of thumb is to use == for almost everything, and reserve is specifically for identity checks like x is None, since None is a singleton and there’s exactly one object to ever compare against.”

Common Mistakes

  • Using is to compare values instead of == — see 7.14 Common Mistakes for why this can silently appear correct due to CPython’s small-integer caching and string interning.
  • Writing x == None instead of the idiomatic x is None — is None is both faster and immune to a custom __eq__ on x that might override equality behavior unexpectedly.

7.9 Operator Precedence and Expression Evaluation

Operator Precedence

What Is It?

The order in which operators are applied when an expression mixes several of them, without any parentheses to force a specific order.

Why Does It Matter?

2 + 3 * 4 is 14, not 20, because * binds tighter than + — getting this wrong silently produces a wrong (but valid-looking) result.

Precedence Table

Highest precedence at top — evaluated first:

PrecedenceOperatorsCategory
Highest()Parentheses — grouping
**Exponentiation
~x, +x, -xUnary invert/plus/minus
* / // %Multiplication, division, modulus
+ -Addition, subtraction
<< >>Bitwise shifts
&Bitwise AND
^ |Bitwise XOR, OR
== != < > <= >=Comparisons
notLogical NOT
andLogical AND
LowestorLogical OR

The practical takeaway: arithmetic > comparisons > logical, so mixed expressions read naturally without needing parentheses everywhere — but add them anyway when in doubt (see 7.17 Best Practices).

Associativity

What Is It?

When operators of the same precedence appear together, associativity decides the grouping direction. Most operators are left-associative (evaluated left to right); ** is right-associative.

>>> 10 - 2 - 3       # left-associative: (10 - 2) - 3
5
>>> 2 ** 3 ** 2       # right-associative: 2 ** (3 ** 2)
512

Parentheses

Explicit parentheses always override default precedence and associativity — the clearest way to guarantee a specific evaluation order, and often better for readability even when not strictly required.

>>> 2 + 3 * 4       # * happens first by default
14
>>> (2 + 3) * 4     # parentheses force + first
20

Expression Evaluation

Simple Expressions

A single operator applied to its operands, producing one value.

>>> 5 + 3
8

Complex Expressions

Multiple operators combined — Python resolves them via precedence, then evaluates left to right within each precedence level, respecting associativity.

>>> x = 5
>>> result = (x + 3) * 2 - 1
>>> result
15

Nested Expressions

Expressions can contain other expressions as operands, to any depth — Python resolves the innermost parentheses first, exactly like standard mathematical notation.

>>> ((2 + 3) * (4 - 1)) ** 2
225

Quick Interview Answer

“Python resolves a mixed expression by precedence first — roughly arithmetic, then bitwise, then comparisons, then logical, from highest to lowest — then by associativity when operators share the same precedence: most are left-associative, but ** is right-associative, so 2 ** 3 ** 2 is 2 ** (3 ** 2) = 512. Parentheses always override both rules and are the clearest way to make evaluation order explicit, which matters most in mixed comparison/logical expressions where the default grouping is easy to misread.”

Common Mistakes

  • Assuming and/or bind tighter than comparisons — comparisons are actually evaluated first, so 2 + 3 == 5 and 1 < 2 groups as ((2+3) == 5) and (1 < 2).
  • Forgetting ** is right-associative, misreading 2 ** 3 ** 2 as (2 ** 3) ** 2 (64) instead of the correct 2 ** (3 ** 2) (512).
  • Nesting nested parentheses so deeply that the expression becomes harder to read than the precedence issue it was meant to avoid.

7.10 Boolean Evaluation

What Is It?

Python treats every value as either “truthy” or “falsy” in a boolean context (if x:, while x:, bool(x)) — not just actual True/False values.

Truthy Values

Most values are truthy by default: non-zero numbers, non-empty strings/lists/dicts, and any object that doesn’t specifically define itself as falsy.

>>> bool(1), bool("a"), bool([1])
(True, True, True)

Falsy Values

A specific, memorizable set of values are falsy: 0, 0.0, "" (empty string), [] (empty list), {} (empty dict), set() (empty set), and None.

>>> bool(0), bool(""), bool([]), bool(None)
(False, False, False, False)

bool() Conversion

Explicitly converts any value to True or False using the truthy/falsy rules above — useful for normalizing a value before storing or comparing it as a boolean.

>>> bool(0), bool(1), bool(""), bool("a"), bool([]), bool([1]), bool(None)
(False, True, False, True, False, True, False)

Quick Interview Answer

“Every value in Python is truthy or falsy in a boolean context, not just literal True/False. The falsy set is small and memorizable: 0, 0.0, '', [], {}, set(), and None — everything else is truthy by default. This is why if my_list: is the idiomatic way to check a list is non-empty instead of if len(my_list) > 0:, and why if x: behaves differently from if x is True: for any value that isn’t literally the boolean True.”

Common Mistakes

  • Writing if len(my_list) > 0: instead of the more idiomatic if my_list: — both work, but the latter is the Pythonic convention and reads more naturally.
  • Confusing if x: (truthy check) with if x is True: (identity check against the literal True) — the second rejects any truthy value that isn’t literally True, like 1 or "yes".
  • Forgetting 0 and 0.0 are falsy when using a numeric value as a flag — a legitimate zero value can accidentally be treated as “unset.”

7.11 Conditional (Ternary) Operator

Syntax

What Is It?

A one-line if/else that produces a value rather than executing a statement block: value_if_true if condition else value_if_false.

Why Is It Used?

For simple two-way choices, it’s more compact than a full if/else block.

status = "adult" if age >= 18 else "minor"
# equivalent to:
# if age >= 18:
#     status = "adult"
# else:
#     status = "minor"

Examples

>>> age = 20
>>> status = "adult" if age >= 18 else "minor"
>>> status
'adult'

Nested Ternary

Chaining ternaries handles more than two outcomes, but readability drops fast past one level of nesting — prefer a full if/elif/else for anything more complex than this.

>>> score = 75
>>> grade = "A" if score >= 90 else "B" if score >= 70 else "C"
>>> grade
'B'

Quick Interview Answer

“The conditional expression value_if_true if condition else value_if_false is Python’s ternary operator — it produces a value rather than executing a statement block, unlike a full if/else. It’s ideal for simple two-way choices assigned to a variable; chaining it for more than two outcomes technically works but readability drops fast, and a full if/elif/else block becomes the better choice past one level of nesting.”

Common Mistakes

  • Nesting ternaries two or more levels deep — technically valid, but forces a reader to mentally trace multiple branches on one line.
  • Using a ternary purely to save lines when the condition or branches are already complex expressions — readability should win over compactness here (see 7.17 Best Practices).

7.12 Chained Comparisons

Syntax

Python allows writing a < b < c directly, unlike languages where you’d need (a < b) and (b < c) explicitly.

>>> x = 5
>>> 1 < x < 10
True

Evaluation

What Happens Internally?

a < b < c is evaluated as (a < b) and (b < c), with b evaluated only once even though it appears twice logically — important if b were an expensive expression or had side effects.

>>> 1 < x < 10 < 20
True

Use Cases

Extremely common for range validation — checking a value falls within bounds in one readable line:

status_code = 404
>>> 400 <= status_code < 500     # is it a client error?
True

Quick Interview Answer

“Chained comparisons let Python express a < b < c directly instead of (a < b) and (b < c). Internally, that’s exactly how it’s evaluated — as an implicit and between each adjacent pair — with one subtlety: the shared middle value (b) is only evaluated once, which matters if it’s an expensive call or has side effects. The most common real use is range validation, like 400 <= status_code < 500 to check an HTTP status code falls in the client-error range.”

Common Mistakes

  • Believing a < b < c evaluates b twice, once for each comparison — it’s evaluated exactly once and reused.
  • Writing a < b and c < d when the intent was actually a < b < c < d — different chains produce very different logic, easy to typo under time pressure.
  • Not realizing chained comparisons work with any comparison operator, not just < — a == b == c and mixed chains like a < b == c are both valid.

7.13 Operators with Different Data Types

Several operators behave completely differently depending on the operand type — + means numeric addition for numbers, but concatenation for sequences. Knowing these per-type behaviors avoids surprises.

Numbers

>>> 5 + 3
8

Strings

+ concatenates; * repeats.

>>> "a" + "b"
'ab'
>>> "ab" * 3
'ababab'

Lists

+ concatenates two lists into a new one; * repeats a list’s elements. See 5.4 Sequence Data Types.

>>> [1, 2] + [3, 4]
[1, 2, 3, 4]
>>> [1, 2] * 2
[1, 2, 1, 2]

Tuples

Same + and * behavior as lists, but the result is always a new tuple, since tuples are immutable — see 5.9 Mutable vs Immutable Types.

>>> (1, 2) + (3, 4)
(1, 2, 3, 4)

Sets

Sets repurpose | and & for union and intersection respectively — a different meaning than their bitwise use on integers (see 7.6 Bitwise Operators), but a natural fit conceptually. See 5.6 Set Data Types.

>>> {1, 2} | {2, 3}     # union
{1, 2, 3}
>>> {1, 2} & {2, 3}     # intersection
{2}

Dictionaries

Python 3.9+ supports | for merging two dicts directly; the ** unpacking trick works on all Python 3 versions. See 5.5 Mapping Data Type.

>>> d1, d2 = {"a": 1}, {"b": 2}
>>> {**d1, **d2}      # merge via unpacking (works on any Python 3)
{'a': 1, 'b': 2}

Quick Interview Answer

“The same operator symbol can mean very different things depending on the operand type: + is numeric addition for numbers but concatenation for strings, lists, and tuples; * is multiplication for numbers but repetition for sequences. Sets repurpose | and & for union and intersection instead of their bitwise meaning on integers. Dicts support | for merging in Python 3.9+, or the {**d1, **d2} unpacking pattern on any Python 3 version — with later keys overriding earlier ones on conflicts.”

Common Mistakes

  • Expecting set1 + set2 to work like list concatenation — sets have no +; use | for union instead.
  • Forgetting {**d1, **d2} merge order matters — keys from d2 override matching keys from d1, not the other way around.
  • Using * to repeat a list of mutable objects (e.g. [[]] * 3) expecting three independent inner lists — it actually creates three references to the same inner list, a classic shared-reference bug (see 6.2 Objects and Variable References).

7.14 Common Mistakes

Using is Instead of ==

What Goes Wrong?

is compares identity, not value — it can appear to work for small integers or short strings, due to CPython’s caching/interning (see 6.7 Memory Optimization), and then mysteriously fail for larger or dynamically-constructed values.

>>> a = 1000
>>> b = int("1000")    # parsed at runtime -- a genuinely separate object
>>> a == b               # correct way to compare VALUES
True
>>> a is b               # unreliable -- False here, though it might look True
False                    # with small cached ints in other examples

Rule: Always use == to compare values. Reserve is strictly for identity checks — most commonly x is None.

Division Mistakes

Confusing / (always float) with // (floored, often int-like) — especially easy to trip over when porting logic from Python 2, where / used to behave like // for two ints.

>>> 5 / 2      # true division
2.5
>>> 5 // 2     # floor division
2

Operator Precedence Issues

Assuming an expression groups the way intended without checking precedence — easy to misread a mixed comparison/logical expression:

>>> result = 2 + 3 == 5 and 1 < 2    # arithmetic and comparisons happen before 'and'
>>> result                            # equivalent to: ((2+3) == 5) and (1 < 2)
True

Quick Interview Answer

“The three recurring operator bugs are: using is instead of == for value comparison, which can silently appear correct thanks to CPython’s small-integer caching and string interning; confusing / (always returns float) with // (floors to a whole number), especially when porting from Python 2; and misjudging operator precedence in a mixed expression, like assuming and binds tighter than == when it’s actually the other way around.”

Common Mistakes

  • Trusting is to “work” in testing with small literal values, then shipping code that breaks on larger or runtime-constructed values — see 7.8 Identity Operators.
  • Assuming // always produces an int — it returns a float if either operand is a float (e.g. 5.0 // 2 is 2.0, not 2).
  • Skipping parentheses in a mixed expression “because it’s obvious,” then having a teammate (or future self) misread it exactly the way the precedence rules predict.

7.15 Performance Considerations

Efficient Expressions

Prefer x in a_set over x in a_list for repeated membership checks on large collections (see 7.7 Membership Operators) — the operator looks identical, but the underlying cost differs by an order of magnitude at scale.

Short-Circuit Benefits

Order and/or conditions so the cheapest or most-likely-to-short-circuit check comes first — e.g. check a cheap flag before calling an expensive function in the same and expression.

# Better: cheap check first, expensive check only runs if needed
if user.is_active and user.has_expensive_permission_check():
    ...

Readability

A marginally “clever” one-liner that saves a few characters is rarely worth the readability cost — favor clear, parenthesized expressions over dense ones (see 7.17 Best Practices).

Quick Interview Answer

“Operator-level performance mostly comes down to two things: choosing the right container for membership tests — set/dict for O(1) average lookups instead of O(n) list scans — and ordering and/or conditions so cheap or likely-to-fail checks run first, letting short-circuit evaluation skip expensive work. Beyond that, a ‘clever’ dense expression rarely outperforms a readable one at runtime; the real cost of cleverness is almost always paid by whoever reads the code next.”

Common Mistakes

  • Micro-optimizing operator choice in code that isn’t actually a bottleneck, at the cost of readability — profile first.
  • Putting an expensive check first in an and/or chain purely because it reads more naturally in that order, missing an easy short-circuit win.

7.16 DevOps Use Cases

Operators are the backbone of every monitoring/validation script — these five patterns cover most of what shows up in practice.

CPU/Disk Threshold Checks

cpu_usage = 87.5
disk_usage = 92.0
CPU_THRESHOLD = 80
DISK_THRESHOLD = 90

>>> if cpu_usage > CPU_THRESHOLD or disk_usage > DISK_THRESHOLD:
...     print("ALERT: Resource threshold exceeded")
ALERT: Resource threshold exceeded

Log Parsing

Comparison and membership operators filter parsed log fields:

log_level = "ERROR"
>>> log_level in {"ERROR", "CRITICAL"}
True

Deployment Validation

Chained comparisons (see 7.12 Chained Comparisons) are the natural way to classify an HTTP status code:

status_code = 404

>>> if 200 <= status_code < 300:
...     result = "Success"
... elif 400 <= status_code < 500:
...     result = "Client Error"
... elif status_code >= 500:
...     result = "Server Error"
...
>>> result
'Client Error'

AWS Resource Checks

instance_state = "running"
>>> is_healthy = instance_state == "running"
>>> is_healthy
True

Health Monitoring

Combining several conditions with and to define “healthy” as a single boolean, benefiting from short-circuit evaluation (see 7.5 Logical Operators) to skip expensive checks once an earlier one already fails:

service_up = True
response_time = 150

>>> is_service_healthy = service_up and response_time < 200
>>> is_service_healthy
True

Quick Interview Answer

“In DevOps scripts, operators show up constantly in a handful of recurring shapes: comparison operators against thresholds for CPU/disk alerts, in against a set of log levels for log filtering, chained comparisons to classify HTTP status codes into success/client-error/server-error bands, == for AWS resource state checks like instance_state == 'running', and and chains combining multiple boolean health signals into one overall health flag — where short-circuit evaluation naturally skips expensive checks once an earlier one has already failed.”

Common Mistakes

  • Comparing an AWS resource state with is instead of == — string identity isn’t guaranteed across API responses, only value equality is reliable.
  • Writing a chain of separate if statements for status-code classification instead of elif with chained comparisons — more error-prone and harder to keep mutually exclusive.
  • Combining unrelated health checks into one dense and expression instead of naming intermediate booleans, covered in 7.17 Best Practices.

7.17 Best Practices

Readable Expressions

Favor clarity over compactness — an expression that takes an extra half-second to parse mentally, multiplied across every future reader, costs far more than the few characters saved.

Avoid Complex Conditions

Break a long chain of and/or into named intermediate booleans — it documents intent and makes each piece independently testable.

# Harder to read at a glance
if cpu > 80 and mem > 90 and not maintenance_mode and disk < 95:
    alert()

# Clearer
resource_critical = cpu > 80 and mem > 90
disk_ok = disk < 95
if resource_critical and disk_ok and not maintenance_mode:
    alert()

Use Parentheses

Even where precedence rules technically make them unnecessary, parentheses make grouping explicit for anyone reading the code without the precedence table memorized (see 7.9 Operator Precedence and Expression Evaluation).

Quick Interview Answer

“The core habits are: favor readable expressions over compact ones, break long and/or chains into named intermediate booleans so intent is documented and each piece is independently testable, and add parentheses even where precedence rules make them technically optional — nobody should have to hold the full precedence table in their head to understand a conditional.”

Common Mistakes

  • Treating a single named boolean as unnecessary overhead for “just one extra condition” — it pays off the moment the condition changes or needs debugging.
  • Adding parentheses inconsistently, making some expressions look ambiguous by contrast with others nearby that have them.

7.18 Interview Questions

Conceptual Questions

  • What is the difference between == and is?
  • What is short-circuit evaluation, and why does it matter?
  • Explain operator precedence with an example that would give a wrong result without it.
  • What’s the difference between / and // in Python 3?
  • How do bitwise operators differ in meaning between integers and sets?

Scenario-Based Questions

  • A status-code check using == works for exact matches, but a range check is needed instead — how would you rewrite it using chained comparison?
  • Two health-check conditions are combined with and, but the second is expensive to compute — how would you order them for best performance?
  • Two config dictionaries need merging, where the second should override the first on conflicts — how do you do that in one line?

Coding Questions

  • Write a function using bitwise operators to check whether a given permission flag is set in a combined permissions integer.
  • Write a one-line ternary expression to classify a number as 'positive', 'negative', or 'zero'.
  • Given a list of HTTP status codes, use membership/comparison operators to count how many are in the 4xx range.

Quick Interview Answer

“The core operator questions test whether the difference between identity and value (is vs ==) is understood, whether short-circuit evaluation is second nature (ordering and/or for both correctness and performance), whether precedence is trusted correctly (arithmetic before comparisons before logical), and the / vs // distinction. The coding questions almost always come down to one of: bit-flag manipulation, a ternary or chained comparison for classification, or a dict-merge one-liner.”

Common Mistakes

  • Answering “is and == are basically the same” without being able to produce a concrete example where they diverge — see 7.8 Identity Operators.
  • Solving the permission-flag coding question with an if/elif chain instead of perms & FLAG, missing the point of the question.
  • Forgetting that {**d1, **d2} (or d1 | d2 on 3.9+) lets the second dict’s keys win on conflict — getting the override direction backwards.

7.19 Hands-on Exercises & Mini Projects

CPU Usage Calculator

Build a function that takes total CPU time used and elapsed time, and returns the percentage utilization using arithmetic operators.

def cpu_utilization(used_seconds, elapsed_seconds):
    return (used_seconds / elapsed_seconds) * 100

>>> cpu_utilization(45, 60)
75.0

Disk Usage Alert

Build a function that returns an alert message using comparison operators against a configurable threshold.

def check_disk(percent_used, threshold=90):
    return "ALERT" if percent_used > threshold else "OK"

>>> check_disk(95)
'ALERT'

Status Code Validator

Build a function using chained comparisons to classify any HTTP status code into its category.

def classify_status(code):
    if 200 <= code < 300:
        return "Success"
    elif 300 <= code < 400:
        return "Redirect"
    elif 400 <= code < 500:
        return "Client Error"
    else:
        return "Server Error"

>>> classify_status(301)
'Redirect'

Service Health Checker

Build a function combining logical operators to determine overall service health from multiple boolean signals.

def is_healthy(is_running, response_time_ms, error_rate):
    return is_running and response_time_ms < 500 and error_rate < 0.05

>>> is_healthy(True, 230, 0.01)
True
>>> is_healthy(True, 800, 0.01)
False

Quick Interview Answer

“These exercises reinforce the chapter’s core ideas hands-on: the CPU calculator exercises arithmetic (/ and *), the disk alert exercises comparison plus the ternary operator, the status validator exercises chained comparisons across an if/elif ladder, and the health checker exercises and-chained logical operators with short-circuit evaluation skipping later checks once an earlier one already fails.”

Common Mistakes

  • Using / where // was intended (or vice versa) in the CPU calculator, producing a fractional result where a whole number was expected, or a truncated one where precision mattered.
  • Writing the status validator as separate if statements instead of if/elif, letting multiple branches match and return inconsistent results.
  • Ordering the health checker’s conditions without considering short-circuit performance — putting an expensive check before a cheap one that’s more likely to fail first.

Operators: Chapter Practice

Mini lab

Compute how many full batches of 10 fit into 23 records and how many remain.

Expected output: 2 and then 3

HintUse the chapter examples to choose the operation, then run your version before opening the solution.
Show solution
print(23 // 10)
print(23 % 10)

Knowledge check

How do / and // differ?

Check your answer/ performs true division; // performs floor division.

Finish the chapter

  • Explain each operation without reading the lesson.
  • Change the sample input and predict the output before running it.
  • Revisit the chapter if the result differs from your prediction.

Back to this chapter · Choose the next chapter

8.1 Introduction to Type Conversion

Python Type Conversion
flowchart TD TC["Type Conversion"] TC --> IM["Implicit\nPython does it automatically"] TC --> EX["Explicit\nyou call the target type yourself"] EX --> NU["Numeric\nint float complex bool"] EX --> ST["String\nstr"] EX --> CO["Collection\nlist tuple set dict"] EX --> BI["Binary\nbytes bytearray memoryview"]

Implicit conversion covers a small, safe set of cases; explicit conversion is everything else this chapter covers.

What Is Type Conversion?

What Is It?

Converting a value from one data type to another — a string "42" becoming the integer 42, for example.

Why Does It Matter?

Data constantly arrives in the “wrong” type for what needs to be done with it (user input is always str, JSON numbers might arrive as str, and so on), so converting reliably is a core everyday skill.

>>> int("42") + 8
50

Why Type Conversion Is Required

Operations are type-specific — arithmetic can’t be done on a string, and a number can’t be concatenated to text, without first converting one side. Type conversion bridges that gap safely and explicitly.

Real-World Examples

  • Converting input() text to a number before doing math on it
  • Converting a config value read as a string ("true") into an actual bool
  • Converting a list of numbers into a set to deduplicate them
  • Converting an object to str for logging or display

Type Conversion in DevOps

Environment variables, CLI arguments, and parsed config/JSON/CSV values all arrive as str by default — correctly converting them to the right type (int, float, bool) is one of the most common sources of bugs in automation scripts, expanded fully in 8.13 Type Conversion in DevOps.

Already Covered Elsewhere

This chapter goes deep on conversion specifically. The groundwork it builds on lives elsewhere:

Dynamic Typing and Conversion

Since a variable’s type is just whatever value it currently holds, “converting” a variable really means creating a new value of the target type and reassigning the variable to it — not modifying the original value in place.

>>> x = "42"
>>> x = int(x)      # x now refers to a completely different, new int object
>>> type(x)
<class 'int'>

Type Compatibility

Not every type can convert to every other type meaningfully — int("abc") fails because "abc" isn’t a valid number, while int("42") succeeds. 8.9 Type Conversion Errors covers exactly which failures to expect and why.

Quick Interview Answer

“Type conversion is turning a value of one type into an equivalent value of another type — Python does this either implicitly, for a small set of always-safe numeric promotions, or explicitly, by calling the target type as a function like int(x) or str(x). Because Python variables are just names bound to objects, ‘converting’ a variable never mutates the original object — it creates a brand-new object of the target type and rebinds the name to it.”

Common Mistakes

  • Assuming a conversion mutates the original value in place — it always creates a new object; the source value is untouched and still independently referenced if anything else points to it.
  • Assuming any two types can convert into each other — conversion only works where a meaningful mapping exists, and fails loudly otherwise (see 8.9 Type Conversion Errors).

8.2 Implicit Type Conversion

flowchart LR subgraph Implicit["Implicit Conversion — Python does it automatically"] I1["int: 1"] --> IR["1 + 2.5 -> 3.5 (float)"] I2["float: 2.5"] --> IR end subgraph Explicit["Explicit Conversion (Casting) — you request it directly"] E1["str: \"42\""] --> E2["int(...)"] --> E3["int: 42"] end

Implicit conversion happens automatically; explicit conversion is requested directly.

Definition

What Is It?

Python automatically converting one type to another in certain mixed-type expressions, without being asked.

Why Is It Used?

For conversions that are always safe and lossless, requiring the programmer to convert manually every time would just be repetitive noise.

Automatic Conversion Rules

The main rule: when int and float are mixed in an arithmetic expression, the int is automatically promoted to float, since float can represent every int value (within precision limits) but not vice versa.

Examples

>>> 1 + 2.5            # int automatically promoted to float
3.5
>>> type(1 + 2.5)
<class 'float'>
>>> True + 1            # bool is a subclass of int -- promotes to int
2

Advantages

  • No boilerplate conversion code needed for safe, common cases
  • Prevents accidental precision loss (never silently truncates int from a float)

Limitations

  • Only works for a small set of built-in, well-defined promotions (mainly numeric)
  • Does not apply between fundamentally incompatible types (str and int do not implicitly combine)
>>> "5" + 3       # NOT automatic -- str and int don't implicitly combine
Traceback (most recent call last):
TypeError: can only concatenate str (not "int") to str

Quick Interview Answer

“Implicit conversion is Python automatically promoting one type to another in a mixed-type expression, without an explicit call — the main real-world case is int promoting to float in mixed arithmetic, since float can safely represent every int value. It’s deliberately narrow: it only covers well-defined, lossless numeric promotions (including bool, since it’s a subclass of int), and it never bridges fundamentally incompatible types like str and int — \"5\" + 3 raises a TypeError rather than silently converting either side.”

Common Mistakes

  • Expecting "5" + 3 to work the way some other dynamically typed languages implicitly coerce strings and numbers — Python raises a TypeError instead.
  • Assuming implicit conversion happens for any “compatible-looking” types — it’s limited to the specific numeric promotions Python defines, not a general rule.

8.3 Explicit Type Conversion (Type Casting)

Definition

What Is It?

Manually requesting a conversion by calling the target type as a function — int(x), str(x), float(x), and so on.

Why Is It Called “Casting”?

Borrowed terminology from languages like C/Java, where a value is explicitly “cast” to a different type.

>>> int("42")
42

Why Explicit Conversion Is Needed

Whenever a conversion isn’t automatic (str to int/float, list to set, and so on) or could lose information (float to int), Python requires an explicit opt-in — this makes the possibility of data loss or an invalid conversion visible in the code, rather than silent. See 8.9 Type Conversion Errors for what happens when the value can’t actually be converted.

Common Use Cases

  • Converting input() text to a number for arithmetic
  • Converting a parsed JSON/CSV string field to its real type
  • Converting a list to a set for deduplication or fast membership testing
  • Converting any object to str for logging, printing, or display

Quick Interview Answer

“Explicit conversion — type casting — is calling the target type directly as a function, like int(x) or str(x), to manually request a conversion Python won’t perform on its own. It’s required for any conversion that isn’t a safe, automatic numeric promotion, or that could lose information, like float to int — Python deliberately makes those visible in the code rather than doing them silently.”

Common Mistakes

  • Reaching for a manual string-parsing workaround instead of the appropriate built-in cast — int(x), float(x), and friends already handle the common cases correctly.
  • Casting a value without first confirming it’s convertible, and letting an unhandled exception crash the program — see 8.10 Safe Type Conversion.

8.4 Numeric Type Conversion

The numeric types themselves are covered in 5.3 Numeric Data Types — this section focuses specifically on the rules for converting into each one.

int()

Converts a string of digits, or truncates a float toward zero (discarding the fractional part — it does not round). Also converts bool (True → 1, False → 0).

>>> int("42")
42
>>> int(3.9)      # truncates -- NOT rounded
3
>>> int(True)
1

float()

Converts a numeric string or an int to a floating-point value.

>>> float("3.14")
3.14
>>> float(5)
5.0

complex()

Converts a number (or a string in a+bj form) into a complex number with real and imaginary parts.

>>> complex(2)
(2+0j)
>>> complex("3+4j")
(3+4j)

bool()

Converts any value to True or False using Python’s truthy/falsy rules — see 8.8 Boolean Conversion and Type Checking.

>>> bool(0), bool(1), bool(""), bool("x")
(False, True, False, True)

Conversion Rules

flowchart LR A["3.99"] --> B["int(...)"] --> C["3"]

int() discards the .99 fractional part — it does not round. Use int(round(3.99)) → 4 if rounding is what’s actually needed.

The most important rule to internalize: int(float_value) always rounds toward zero, never to the nearest whole number. Use round() first if rounding is actually the goal:

>>> int(3.9), int(-3.9)      # both truncate toward zero
(3, -3)
>>> int(round(3.9))           # round first, then convert
4

Quick Interview Answer

“int() parses a digit string or truncates a float toward zero — critically, it does not round, so int(3.9) is 3, and int(-3.9) is -3, not -4. float() parses a numeric string or widens an int. complex() builds a complex number from a real number or an a+bj string. bool() applies the truthy/falsy rules explicitly. When rounding to the nearest whole number is actually the goal, round() first and int() second: int(round(3.9)) is 4.”

Common Mistakes

  • Expecting int(3.9) to round to 4 — it truncates toward zero, giving 3; use round() first for actual rounding.
  • Forgetting int(-3.9) truncates to -3 (toward zero), not -4 (toward negative infinity, which is what // does) — the two are easy to conflate.
  • Trying int("3.14") directly — it raises ValueError, since int() expects a string that’s already a whole number; go through float() first (see 8.15 Common Mistakes).

8.5 String Conversion

str()

Converts virtually any object into its human-readable string representation — used constantly for logging, printing, and building messages.

>>> str(42)
'42'

Converting Numbers to Strings

>>> str(42), str(3.14), str(True)
('42', '3.14', 'True')

Converting Collections to Strings

str() on a list/dict/tuple produces a debug-friendly text representation matching how it would be typed as a literal — useful for logging, but not the format to show an end user (for that, format the contents explicitly).

>>> str([1, 2, 3])
'[1, 2, 3]'
>>> str({"a": 1})
"{'a': 1}"

Quick Interview Answer

“str() converts virtually any object into its human-readable string form — for numbers that’s the obvious text representation, and for a collection it produces the same text you’d type to construct it as a literal (str([1, 2, 3]) is '[1, 2, 3]'). That collection representation is meant for debugging and logging, not end-user display — showing a raw str(some_dict) to a user is a common shortcut that reads poorly; format the contents deliberately instead.”

Common Mistakes

  • Using str(some_list) or str(some_dict) directly in user-facing output instead of formatting the contents explicitly — technically works, but reads like debug output.
  • Forgetting str(True) is 'True' (capitalized), not 'true' — a frequent source of bugs when building JSON-like text manually instead of using the json module.

8.6 Collection Type Conversion

flowchart TD STR["str"] <--> INT["int"] STR <--> FLOAT["float"] INT <--> FLOAT STR <--> COLL["list / tuple / set"] BOOL["bool"] --> INT BOOL --> FLOAT

Common conversion paths between core Python types.

list()

Converts any iterable (string, tuple, set, range, dict keys, …) into a list.

>>> list("abc")
['a', 'b', 'c']

tuple()

Converts any iterable into an immutable tuple — useful when a fixed, hashable version of a collection is needed. See 5.9 Mutable vs Immutable Types.

>>> tuple([1, 2, 3])
(1, 2, 3)

set()

Converts any iterable into a set, automatically removing duplicates — the standard one-line deduplication idiom.

>>> set([1, 1, 2, 3])
{1, 2, 3}

frozenset()

The immutable counterpart to set() — produces a set that can itself be used as a dict key or set member. See 5.6 Set Data Types.

>>> frozenset([1, 2, 3])
frozenset({1, 2, 3})

dict()

Converts an iterable of key-value pairs (like a list of 2-tuples) into a dict.

>>> dict([("a", 1), ("b", 2)])
{'a': 1, 'b': 2}

range()

Not exactly a “conversion” target (nothing converts into range), but frequently converted from — generates numbers lazily, and list(range(...)) materializes them:

>>> list(range(5))
[0, 1, 2, 3, 4]

Quick Interview Answer

“list(), tuple(), set(), frozenset(), and dict() all accept any iterable — a string, another collection, a range, dict keys — and build a new collection of their own type from it. set() and frozenset() are the standard deduplication idiom, dropping duplicates automatically. dict() specifically expects an iterable of key-value pairs, like a list of 2-tuples. range is unusual: nothing converts into it, but it’s frequently converted from, via list(range(...)), to materialize its lazily generated numbers.”

Common Mistakes

  • Passing a plain list of values (not pairs) to dict() expecting it to work — it requires an iterable of key-value pairs, like [(k1, v1), (k2, v2)], and raises otherwise.
  • Materializing a large range() into a list() unnecessarily, losing its lazy memory efficiency for no reason (see 5.4 Sequence Data Types).
  • Converting a list with unhashable elements (like nested lists) to a set, hitting TypeError: unhashable type.

8.7 Binary Type Conversion

bytes, bytearray, and memoryview themselves are covered in 5.7 Binary Data Types — this section focuses on the conversion rules for constructing each from other types, essential whenever Python touches files, sockets, or non-text data.

bytes()

Converts a list of integers (0–255), or a str with an explicit encoding, into an immutable bytes object.

>>> bytes([65, 66, 67])
b'ABC'
>>> bytes("hi", "utf-8")
b'hi'

bytearray()

Same conversion rules as bytes(), but produces a mutable result that can be modified in place afterward.

>>> bytearray("hi", "utf-8")
bytearray(b'hi')

memoryview()

Wraps an existing bytes/bytearray object as a zero-copy view — converting it back to a list of integers shows the underlying byte values without duplicating the buffer.

>>> list(memoryview(b"hi"))
[104, 105]

Quick Interview Answer

“bytes() builds an immutable byte sequence either from a list of integers in the 0–255 range, or from a str given an explicit encoding — Python never guesses an encoding, it must be stated. bytearray() follows the same construction rules but produces a mutable result. memoryview() is different in kind: it doesn’t copy data at all, just wraps an existing bytes/bytearray buffer, which matters for processing large binary data without duplicating it in memory.”

Common Mistakes

  • Calling bytes("hi") without an encoding argument — unlike some conversions, str to bytes requires an explicit encoding; Python won’t assume one.
  • Expecting memoryview() to copy data like the other conversions do — it’s a zero-copy view onto the original buffer, not a new independent object.

8.8 Boolean Conversion and Type Checking

Boolean Conversion

The truthy/falsy rules themselves are covered from the operator angle in 7.10 Boolean Evaluation — bool() is simply the explicit-casting form of that same logic, worth repeating here specifically as a conversion.

>>> bool(1), bool("x"), bool([1])          # truthy
(True, True, True)
>>> bool(0), bool(""), bool([]), bool(None)   # falsy
(False, False, False, False)

bool() is the single function underlying every truthy/falsy check Python makes internally (in if, while, and, or) — calling it directly makes that implicit check explicit and visible.

Checking Data Types

Before or after converting, it’s often necessary to verify a value’s actual type — covered in depth in 5.10 Type Checking, summarized here for this chapter’s context.

type()

Returns a value’s exact type — useful for confirming a conversion succeeded and produced what was expected.

>>> x = int("42")
>>> type(x)
<class 'int'>

isinstance()

Checks whether a value is an instance of a type (including subclasses) — the preferred way to validate a value’s type before attempting to convert or use it.

>>> isinstance(42, int)
True
>>> isinstance("42", int)      # a string is NOT an int, even if it looks numeric
False

Quick Interview Answer

“bool() is the explicit-casting counterpart to Python’s truthy/falsy rules — the same logic if/while/and/or apply implicitly, made visible as an actual function call. After converting, type() confirms the exact resulting type, while isinstance() is the preferred check before converting or using a value, since it correctly accounts for subclasses — isinstance(True, int) is True because bool subclasses int, which type(x) == int would miss.”

Common Mistakes

  • Assuming a string that looks like a boolean converts sensibly with bool() — see 8.15 Common Mistakes for why bool("False") is True.
  • Using type(x) == SomeClass instead of isinstance(x, SomeClass) before deciding how to convert a value — breaks for subclass instances that isinstance() would correctly accept.

8.9 Type Conversion Errors

Three exception types account for nearly every conversion failure — recognizing them immediately tells you what went wrong.

ValueError

Raised when the type is correct but the actual value can’t be parsed into the target type — e.g. a string that isn’t a valid number.

>>> int("abc")
Traceback (most recent call last):
ValueError: invalid literal for int() with base 10: 'abc'

TypeError

Raised when the conversion doesn’t even make sense for that type at all — e.g. trying to convert None or a list directly into an int.

>>> int(None)
Traceback (most recent call last):
TypeError: int() argument must be a string, a bytes-like object or a real number, not 'NoneType'

OverflowError

Raised when a numeric result is too large to be represented — rare with Python’s arbitrary-precision ints, but real for float, which has finite range.

>>> float(2) ** 10000
Traceback (most recent call last):
OverflowError: (34, 'Numerical result out of range')

Common Causes

  • Converting user input without validating it first (see 8.10 Safe Type Conversion)
  • Assuming an API/CSV/JSON field is always the type expected
  • Trying to parse a decimal string directly with int() instead of going through float() first

Quick Interview Answer

“Three exceptions cover almost every conversion failure: ValueError, when the type is right but the value can’t be parsed — like int('abc'); TypeError, when the conversion doesn’t make sense for that type at all — like int(None); and OverflowError, when a numeric result is too large to represent, which is rare for Python’s arbitrary-precision int but real for float’s finite range. The most common real-world trigger is converting untrusted input — user input, an API field, a CSV cell — without validating it first.”

Common Mistakes

  • Catching a bare except: instead of the specific exception type, hiding bugs unrelated to the conversion itself.
  • Not distinguishing ValueError from TypeError when handling a conversion — both can happen from the same call site depending on the input, and often both need catching together (see 8.10 Safe Type Conversion).

8.10 Safe Type Conversion

flowchart TD V["Untrusted value\n(e.g. user input, env var)"] --> T["try:\nint(value)"] T -->|"valid"| S["Success:\nreturn converted value"] T -->|"invalid"| E["except (ValueError, TypeError):\nreturn default / log / re-prompt"]

Wrapping a conversion in try/except lets the program degrade gracefully instead of crashing.

Using try/except

What Is It?

Wrapping a conversion call in try/except so an invalid input produces a controlled fallback instead of crashing the whole program.

def safe_int(value, default=0):
    try:
        return int(value)
    except (ValueError, TypeError):
        return default

>>> safe_int("42")
42
>>> safe_int("abc")
0
>>> safe_int(None)
0

Input Validation

Validating before converting (e.g. checking str.isdigit() first) can avoid the exception path entirely for simple cases, though try/except remains the more general and robust approach for anything beyond plain digit strings.

Error Handling

Beyond just returning a default, a production script should typically also log the invalid input it encountered, so bad data upstream gets noticed and fixed rather than silently swallowed forever.

Quick Interview Answer

“The standard pattern for safe conversion is wrapping the call in try/except (ValueError, TypeError) and returning a sensible default or re-prompting, rather than letting the exception propagate and crash the program. Pre-validating with something like str.isdigit() can skip the exception path for simple digit strings, but try/except is the more general approach — it correctly handles every failure mode, not just the ones a hand-written validator anticipated. In production code, a caught conversion failure is usually also worth logging, so bad upstream data gets noticed rather than silently defaulted away forever.”

Common Mistakes

  • Catching Exception broadly instead of the specific (ValueError, TypeError) pair, masking unrelated bugs.
  • Returning a default silently with no logging, letting systematically bad upstream data go unnoticed indefinitely.
  • Relying only on str.isdigit() pre-validation and skipping try/except entirely — it doesn’t handle negative numbers, decimals, or non-string inputs like None.

8.11 Memory and Performance

Object Creation During Conversion

Every conversion call (int(x), str(x), …) creates a brand-new object — it never modifies the original in place, consistent with immutability (see 5.9 Mutable vs Immutable Types). The original value remains unchanged and independently referenced until nothing points to it anymore.

Performance Considerations

Converting inside a tight loop repeats the allocation cost every iteration — if the same value is converted many times, convert it once and reuse the result instead.

# Wasteful: re-converts the same value every iteration
for _ in range(1000):
    threshold = int(config["threshold"])

# Better: convert once, reuse
threshold = int(config["threshold"])
for _ in range(1000):
    ...    # use threshold directly

Memory Usage

Different types have different memory footprints for logically similar data — e.g. a tuple is generally more compact than an equivalent list. Converting to a heavier type unnecessarily wastes memory at scale.

Quick Interview Answer

“Every conversion allocates a brand-new object — nothing is ever modified in place, since Python’s conversions produce independent values consistent with how immutability works. The practical performance implication is to convert once and reuse the result, rather than re-converting the same value on every loop iteration, which repeats the allocation cost for no benefit. It’s also worth choosing the lightest type that actually fits the use case — converting to a heavier structure than needed wastes memory at scale for no functional gain.”

Common Mistakes

  • Re-running a conversion inside a loop body on a value that never changes — an easy, invisible performance cost that grows with iteration count.
  • Converting to a mutable collection (like list) purely out of habit when an immutable, more compact one (tuple) would serve the same purpose.

8.12 Type Conversion in File Handling

Every one of these external data sources hands over strings (or, for JSON, occasionally the wrong type) — converting on the way in is a universal pattern.

Reading User Input

input() always returns str, regardless of what’s typed — explicit conversion is mandatory before any arithmetic.

>>> age_str = input("Enter age: ")     # ALWAYS str, even for '25'
>>> age = int(age_str)
>>> age + 1
26

CSV Data

The csv module reads every field as str, with no automatic type inference — numeric fields must be explicitly converted after reading.

>>> import csv, io
>>> row = next(csv.DictReader(io.StringIO("name,age\nAlice,30")))
>>> row, type(row["age"])
({'name': 'Alice', 'age': '30'}, <class 'str'>)
>>> age = int(row["age"])
>>> age, type(age)
(30, <class 'int'>)

JSON Data

Unlike CSV, json.loads() does infer types automatically for genuine JSON numbers/booleans — but a field written as a JSON string ("8080" instead of 8080) still needs manual conversion.

>>> import json
>>> data = json.loads('{"port": "8080"}')     # port is a JSON string here
>>> type(data["port"])
<class 'str'>
>>> port = int(data["port"])
>>> port, type(port)
(8080, <class 'int'>)

YAML Data

PyYAML’s safe_load() generally infers types well (numbers, booleans, and null all come back as their proper Python types) — but values explicitly quoted in the YAML source still arrive as str and may need conversion, same as JSON.

Quick Interview Answer

“Every external text-based data source needs the same discipline: input() always returns str, no exceptions. The csv module never infers types at all — every field is str regardless of content. json.loads() and YAML’s safe_load() do infer real types for genuine JSON/YAML numbers and booleans, but a value that was written as a quoted string in the source — like \"8080\" instead of 8080 — comes back as str even though it looks numeric, and still needs an explicit conversion after parsing.”

Common Mistakes

  • Assuming csv.DictReader infers numeric columns — it never does; every field is str until explicitly converted.
  • Assuming every JSON number field is automatically an int/float in Python — only true if it was written as a JSON number in the source, not as a quoted string.
  • Comparing a CSV or JSON string field directly against a number without converting first — see 8.13 Type Conversion in DevOps for the exact failure mode this causes.

8.13 Type Conversion in DevOps

Five places type conversion shows up constantly in real infrastructure scripts.

Environment Variables

os.environ always stores values as str — numeric or boolean-looking env vars must be explicitly converted.

>>> import os
>>> os.environ["MAX_RETRIES"] = "5"
>>> retries = int(os.environ.get("MAX_RETRIES", "3"))
>>> retries, type(retries)
(5, <class 'int'>)

Configuration Files

INI-style config files (via configparser) also return every value as str — the same explicit-conversion discipline applies as with environment variables.

AWS API Responses

boto3 generally returns properly-typed Python values (int, bool, datetime) directly from AWS APIs — but any value extracted from a raw JSON blob in your own code still needs the same JSON-conversion care as 8.12 Type Conversion in File Handling.

Log Parsing

Every field extracted from a raw log line (via regex or split()) starts life as str — durations, status codes, and counts all need explicit conversion before comparison or arithmetic.

status_str = "404"
>>> status_code = int(status_str)
>>> 400 <= status_code < 500
True

HTTP Status Codes

A common real mistake: comparing a status code that’s still a str against an int range, which silently never matches instead of raising an error — always confirm the type before a range/comparison check.

>>> "404" == 404      # different types -- always False, no error!
False

Quick Interview Answer

“In DevOps scripts, type conversion shows up in five recurring places: environment variables (os.environ is always str), config files (configparser is also always str), AWS API responses (boto3 returns properly-typed values, but anything pulled from a raw JSON blob still needs manual conversion), log parsing (every extracted field starts as str), and HTTP status code comparisons specifically — comparing a still-str status code against an int doesn’t raise an error, it just silently and permanently evaluates False, which is a genuinely dangerous failure mode because nothing signals it went wrong.”

Common Mistakes

  • Comparing os.environ.get("PORT") == 8080 — the env var is a str, so this is always False no matter its actual value, with no exception raised to flag the bug.
  • Trusting boto3 return types blindly for values that actually originated from a nested raw JSON/string field in the response rather than a native AWS API type.
  • Not converting a log-parsed duration or count before doing arithmetic on it, producing string concatenation instead of addition.

8.14 Best Practices

Choose the Right Data Type

Convert to the type that actually matches how the value will be used — don’t leave a numeric value as str just because that’s how it arrived.

Avoid Unnecessary Conversions

Convert once, as early as possible after receiving external data, and work with the properly-typed value from then on — repeatedly converting the same value back and forth wastes effort and invites bugs. See 8.11 Memory and Performance.

Validate User Input

Never assume input() or an external field will parse cleanly — wrap conversions of untrusted data in try/except (see 8.10 Safe Type Conversion) as standard practice, not an afterthought.

Quick Interview Answer

“Three habits cover most of what matters: convert to the type that actually matches how a value will be used, rather than leaving everything as str; convert once, as early as possible after receiving external data, instead of repeatedly re-converting the same value; and treat validating untrusted input — user input, env vars, API fields — as a default, not something bolted on after a bug report.”

Common Mistakes

  • Leaving a value as str throughout a codebase “because it arrived that way,” pushing the conversion responsibility onto every caller instead of doing it once at the boundary.
  • Adding input validation only after a production incident, instead of treating it as the default for any data crossing a trust boundary.

8.15 Common Mistakes

Invalid Conversions

Trying to int() a decimal-looking string directly fails, because int() expects a string that’s already a whole number — go through float() first if the string might have a decimal point.

>>> int("3.14")
Traceback (most recent call last):
ValueError: invalid literal for int() with base 10: '3.14'
>>> int(float("3.14"))    # correct two-step approach
3

Data Loss

Converting float to int silently discards the fractional part (see 8.4 Numeric Type Conversion) — easy to miss if not paying attention, since no error or warning is raised.

>>> price = 19.99
>>> int(price)      # silently loses the cents -- no error at all
19

Incorrect Boolean Conversion

What Goes Wrong?

Assuming bool("False") or bool("0") returns False, because the text looks like a negative value. In reality, any non-empty string is truthy, regardless of its content.

>>> bool("False")     # non-empty string -- ALWAYS True
True
>>> bool("0")          # also True -- a very common gotcha
True

# Correct way to parse a boolean-looking string:
>>> value = "False"
>>> value.lower() == "true"
False

Quick Interview Answer

“Three conversion bugs come up constantly: calling int() directly on a decimal-looking string, which raises ValueError because int() only accepts strings that are already whole numbers — the fix is int(float(x)); silently losing precision converting float to int, since truncation raises no warning at all; and assuming bool() parses boolean-looking text semantically — bool(\"False\") and bool(\"0\") are both True, because any non-empty string is truthy regardless of content. Parsing an actual boolean-looking string correctly means comparing its lowercased text against \"true\" explicitly, not calling bool() on it.”

Common Mistakes

  • Calling bool(some_string) to parse a config value that’s textually "true"/"false" — it always returns True for any non-empty string; compare the lowercased text instead.
  • Not noticing int(19.99) silently drops the cents — worth an explicit round() first if the intent is rounding rather than truncation.
  • Skipping the float() intermediate step when parsing a decimal string with int(), hitting an avoidable ValueError.

8.16 Interview Questions

Conceptual Questions

  • What’s the difference between implicit and explicit type conversion?
  • Why does int(3.9) return 3 instead of 4?
  • What’s the difference between ValueError and TypeError during conversion?
  • Why is bool("False") True?
  • Why does input() always require explicit conversion for numeric use?

Scenario-Based Questions

  • A script reads a config value "true" from a file and needs it as a real bool — how would you convert it correctly?
  • An API sometimes returns a numeric field as a JSON string instead of a number — how would you defensively handle both cases?
  • A CSV import is crashing on a few malformed rows — how would you make the numeric conversion resilient to bad data?

Coding Questions

  • Write a safe_float() function that returns a default value instead of raising on invalid input.
  • Write a function that correctly parses common boolean-like strings ("true"/"false"/"yes"/"no") into actual bool values.
  • Given a list of mixed-type values, write code that converts only the ones that can be converted to int, skipping the rest.

Quick Interview Answer

“The core conversion questions test whether implicit vs explicit conversion is understood clearly, whether int()’s truncate-toward-zero behavior is known (not rounding), whether ValueError vs TypeError can be distinguished by cause, and whether the bool(\"False\") gotcha is anticipated. The coding questions almost always come down to writing a safe_* wrapper around try/except, or defensively normalizing a value that might arrive as either its real type or a string version of it.”

Common Mistakes

  • Answering “int(3.9) rounds to 4” — it truncates toward zero to 3; conflating truncation with rounding is one of the most common wrong answers here.
  • Writing a boolean-string parser using bool(value) directly instead of comparing lowercased text against "true" — misses the entire point of the question.
  • Solving the “convert only what’s convertible” coding question with a bare except: instead of the specific (ValueError, TypeError) pair.

8.17 Hands-on Exercises

User Input Converter

Build a function that safely converts input() text to int, float, or bool based on a requested target type, with sensible error handling.

def convert_input(value, target_type):
    try:
        if target_type == bool:
            return value.strip().lower() in ("true", "yes", "1")
        return target_type(value)
    except (ValueError, TypeError):
        return None

>>> convert_input("42", int)
42
>>> convert_input("yes", bool)
True
>>> convert_input("abc", int)
None

CSV Parser

Build a function that reads CSV rows and converts specified numeric columns, handling any row with bad data gracefully.

import csv, io

def parse_ages(csv_text):
    reader = csv.DictReader(io.StringIO(csv_text))
    result = []
    for row in reader:
        try:
            row["age"] = int(row["age"])
        except ValueError:
            row["age"] = None
        result.append(row)
    return result

>>> parse_ages("name,age\nAlice,30\nBob,N/A")
[{'name': 'Alice', 'age': 30}, {'name': 'Bob', 'age': None}]

Log Analyzer

Build a function that extracts and converts status codes from log lines, classifying each safely even if a line is malformed.

def classify_status_line(status_str):
    try:
        code = int(status_str)
    except ValueError:
        return "invalid"
    if 200 <= code < 300:
        return "success"
    elif 400 <= code < 500:
        return "client_error"
    elif code >= 500:
        return "server_error"
    return "other"

>>> classify_status_line("404")
'client_error'
>>> classify_status_line("N/A")
'invalid'

Mini Projects

  • Config loader — reads a dict of raw string values and converts each to its declared type (int/float/bool), reporting any that fail.
  • CLI calculator — safely converts two input() values to float and applies a chosen arithmetic operator, handling invalid input without crashing.
  • Environment-variable validator — checks a list of required env vars exist and are convertible to their expected types before a script proceeds.

Quick Interview Answer

“These exercises reinforce the chapter’s core ideas hands-on: the input converter exercises try/except around a dynamic target type plus the boolean-string gotcha, the CSV parser exercises per-row resilience so one bad record doesn’t crash the whole import, and the log analyzer combines safe conversion with chained-comparison classification from 7.12 Chained Comparisons.”

Common Mistakes

  • Letting one malformed CSV row crash the entire parse instead of catching the conversion failure per-row and continuing.
  • Forgetting the boolean branch in convert_input needs its own logic — calling bool(value) directly on a string would always return True for any non-empty input.
  • Not handling the case where target_type itself is invalid or unexpected, instead of assuming callers always pass one of the anticipated types.

Type Conversion: Chapter Practice

Mini lab

Parse a CPU reading from “72.5” and compare it with 70.

Expected output: True

HintUse the chapter examples to choose the operation, then run your version before opening the solution.
Show solution
cpu = float("72.5")
print(cpu >= 70)

Knowledge check

Is bool(“False”) false?

Check your answerNo. Every nonempty string is truthy.

Finish the chapter

  • Explain each operation without reading the lesson.
  • Change the sample input and predict the output before running it.
  • Revisit the chapter if the result differs from your prediction.

Back to this chapter · Choose the next chapter

9.1 Introduction to Strings

Python Strings
flowchart TD S["str"] S --> C["Creation\nquotes, raw, unicode"] S --> I["Indexing & Slicing"] S --> F["Formatting\n% .format f-strings"] S --> M["Methods\ncase, search, split/join"] S --> A["Algorithms\nreverse, palindrome, anagram"] S --> R["Regex\nre module"]

Everything in this chapter builds on one fact: a string never changes after it’s created.

What Is a String?

What Is It?

A string is an ordered, immutable sequence of Unicode characters — Python’s built-in type for representing text. Anything that is “text” in a Python program, from a single letter to an entire log file’s contents, is stored as a str object.

>>> s = "DevOps"
>>> type(s)
<class 'str'>
>>> len(s)
6

Why Does It Matter?

Strings are the universal interface between a program and the outside world — every system boundary (files, networks, users, other programs) communicates in text. Every filename, URL, config value, and piece of user input passes through str, and string-handling correctness (encoding, escaping, formatting) is a common source of production bugs.

Real-World Uses of Strings

  • Parsing configuration files (JSON, YAML, INI)
  • Processing application and server logs
  • Building and validating URLs, emails, and file paths
  • Constructing SQL/NoSQL queries and API request bodies
  • Templating reports, emails, and generated code

Strings in DevOps and AWS

In DevOps and cloud engineering specifically, strings are the interface: CloudWatch logs, Terraform output, kubectl output, ARNs, and CI/CD console logs are all just text that scripts must parse reliably — the focus of 9.12 Strings in DevOps and AWS later in this chapter.

Memory Representation of Strings

Every Python string is a heap-allocated object (PyUnicodeObject in CPython) carrying a reference count, type pointer, cached hash, and a character buffer. A variable holding a string is just a reference to this object — assignment copies a pointer, never the characters themselves (see 6.3 Memory Management: Stack vs Heap).

CPython also interns many literal, identifier-like strings, so equal literals can share a single object:

>>> a = "hello"; b = "hello"
>>> a is b
True     # same interned object -- not guaranteed for runtime-built strings

String Immutability

A string can never be changed in place. Every method that appears to modify a string (replace(), upper(), strip(), …) returns a brand-new object; the original is untouched — the same rule covered generally in 5.9 Mutable vs Immutable Types.

>>> s = "hello"
>>> s.upper()
'HELLO'
>>> s          # unchanged
'hello'

Advantages and Limitations

AdvantagesLimitations
Safe to share across functions/threads (no aliasing bugs)Every “modification” allocates a new object
Usable as dict keys / set members (hashable)Looping with += is O(n²) for large strings — see 9.11 Memory and Performance
Interning saves memory for repeated literalsFixed per-object overhead beyond raw characters
Rich standard-library method setWide Unicode strings use more memory per character

Quick Interview Answer

“A string is an ordered, immutable sequence of Unicode characters — Python’s built-in type for text. Immutability is the defining property: no method ever changes a string in place, every apparent modification returns a brand-new object, and that’s exactly what makes strings safe to share across functions and threads and usable as dict keys. Under the hood, every string is a heap-allocated PyUnicodeObject, and CPython interns many literal strings so identical literals can share one object — though that’s an implementation detail, not something to rely on for correctness.”

Common Mistakes

  • Assuming s.replace(...) or s.upper() mutates s in place — it returns a new string; the call is useless unless the result is captured or reassigned.
  • Comparing strings with is instead of == because interning happened to make a small test case pass — interning isn’t guaranteed for every string, especially ones built at runtime.
  • Treating a string as mutable “because it looks like a list of characters” — indexing works, but s[0] = "x" raises TypeError.

9.2 Creating Strings

Single and Double Quotes

Single and double quotes are functionally identical — the choice is purely stylistic, except when the text itself contains a quote character, in which case using the other kind avoids escaping (see 9.3 Escape Characters).

>>> s = 'Hello'
>>> s = "It's DevOps"      # double quotes avoid escaping the apostrophe
"It's DevOps"

Triple Quotes and Multiline Strings

Triple single or double quotes ('''...''' or """...""") let a string span multiple lines and contain unescaped quote characters. They’re also how Python docstrings are written, and are the standard way to embed SQL queries, HTML/email templates, or usage text directly in code.

>>> query = """
... SELECT * FROM users
... WHERE active = true
... """
>>> print(query)
SELECT * FROM users
WHERE active = true

Empty Strings

An empty string ("") is a valid, zero-length str — often used as an accumulator’s starting value or a default “no value” placeholder. It’s falsy in a boolean context, a common way to check for missing text.

>>> s = ""
>>> len(s)
0
>>> bool(s)      # empty string is falsy
False

Raw Strings (r"...")

A string literal prefixed with r treats backslashes as literal characters instead of escape-sequence introducers — essential for Windows file paths and regex patterns (see 9.10 Regular Expressions with Strings), both of which use backslashes heavily.

>>> path = r"C:\new\test"
>>> path
'C:\\new\\test'

Unicode Strings

Every Python 3 str is Unicode by default — no special prefix needed, unlike Python 2’s u"..." — so accented letters, symbols, and non-Latin scripts just work out of the box.

>>> s = "héllo wörld ünïcödé"
>>> print(s)
héllo wörld ünïcödé

Quick Interview Answer

“Single and double quotes are interchangeable; pick whichever avoids escaping a quote character inside the text. Triple quotes span multiple lines and are also how docstrings are written. A raw string literal (r"...") turns off backslash escaping, which is essential for Windows paths and regex patterns that would otherwise need every backslash doubled. Every Python 3 str is Unicode natively — there’s no separate byte-string-by-default mode like Python 2 had.”

Common Mistakes

  • Forgetting a raw string’s backslash still can’t end the literal with an odd number of trailing backslashes (r"path\" is a syntax error) — Python’s raw-string handling isn’t quite “no escaping at all.”
  • Using a regular string for a Windows path or regex pattern and getting mangled output from unintended escape sequences like \n or \t appearing mid-path.
  • Assuming triple-quoted strings are a distinct type from single/double-quoted ones — they all produce the exact same str object; triple quotes only affect what characters can appear unescaped.

9.3 Escape Characters

The core escape sequences (\n, \t, \\, quotes) were introduced in 4.13 Escape Characters. This is the complete reference, including the octal and hex forms.

EscapeMeaningExample
\nNewline"a\nb" → two lines
\tTab"a\tb" → "a b"
\\Literal backslash"a\\b" → a\b
\'Literal single quote'it\'s' → it's
\"Literal double quote"say \"hi\"" → say "hi"
\rCarriage returnused in some line-ending formats
\bBackspacemoves cursor back one position
\fForm feedpage-break control character
\oooCharacter by octal code\101 → 'A'
\xhhCharacter by hex code\x41 → 'A'
>>> print("\101\102\103")     # octal escapes
ABC
>>> print("\x41\x42\x43")     # hex escapes
ABC

Quick Interview Answer

“Beyond the everyday \n, \t, and quote escapes, Python also supports character-by-code escapes: \ooo for octal and \xhh for hex. Both resolve to a single character at parse time — \x41 and \101 are two different ways of writing the letter 'A'. These are rarely needed for hand-written code but show up when strings are generated programmatically from raw byte or code-point values.”

Common Mistakes

  • Confusing \x41 (a single hex-escaped character) with a literal two-character sequence — it always collapses to one character in the resulting string.
  • Writing an octal escape with a digit \8 or \9 and expecting it to work — octal digits only go up to 7, so those aren’t valid octal escapes.
  • Reaching for raw strings (r"...") when the octal/hex escapes are actually needed — a raw string disables all escape processing, including these.

9.4 String Indexing and Slicing

Indexing and Negative Indexing

Every character in a string has a position. Positive indices count from the left starting at 0; negative indices count from the right starting at -1 — useful for reaching the end of a string without knowing its length.

>>> s = "Programming"
>>> s[0], s[-1]
('P', 'g')

Slicing

Slicing extracts a sub-string using s[start:stop:step]. start is the index to begin at (inclusive), stop is the index to end before (exclusive), and step is the gap between characters taken. Any of the three can be omitted — start defaults to 0, stop defaults to the end of the string, step defaults to 1.

flowchart LR P0["P\n0 / -6"] --- Y1["Y\n1 / -5"] --- T2["T\n2 / -4"] --- H3["H\n3 / -3"] --- O4["O\n4 / -2"] --- N5["N\n5 / -1"]

Positive indices count from the left; negative indices count from the right of the string "PYTHON".

>>> s = "Programming"
>>> s[2:5]     # characters at index 2, 3, 4 -- stop is exclusive
'ogr'
>>> s[:4]      # start defaults to 0
'Prog'
>>> s[4:]      # stop defaults to end of string
'ramming'
>>> s[:]       # both default -- a full copy-like slice
'Programming'

Slicing Never Raises IndexError

Unlike direct indexing (s[50], which raises IndexError if out of range), a slice with an out-of-range start or stop simply clamps to the available characters — or returns an empty string if the range doesn’t overlap the string at all. This makes slicing safe to use with indices that haven’t been bounds-checked.

>>> s = "Programming"
>>> s[2:100]      # stop far beyond the end -- clamps, no error
'ogramming'
>>> s[100:200]    # entirely out of range -- empty string, no error
''

Extended Slicing

Extended slicing adds the step argument. A step greater than 1 skips characters; a negative step walks backward, which is what makes s[::-1] the idiomatic way to reverse a string — see 9.9 String Algorithms.

>>> s = "Programming"
>>> s[::2]      # every 2nd character, left to right
'Pormig'
>>> s[::-1]     # negative step -- walks backward -> reverses the string
'gnimmargorP'
>>> s[5:1:-1]   # explicit start/stop with a negative step
'argo'

Real-world examples — stripping a known suffix/prefix and pulling a fixed-width timestamp out of a log line:

>>> filename = "server_backup_2026.tar.gz"
>>> filename[:-7]      # strip the 7-char '.tar.gz' suffix
'server_backup_2026'

>>> log_line = "2026-07-13T10:22:05Z INFO message here"
>>> log_line[:19]      # fixed-width ISO timestamp
'2026-07-13T10:22:05'

Copying Strings

Because strings are immutable, Python doesn’t need to make a real copy when the whole string is sliced — CPython recognizes s[:] and simply returns the same object, which is safe precisely because that object can never be mutated.

>>> s = "Programming"
>>> s_copy = s[:]
>>> s_copy is s     # CPython reuses the same object -- safe due to immutability
True

Quick Interview Answer

“Indexing uses s[i], with negative indices counting from the end. Slicing, s[start:stop:step], extracts a sub-string — stop is exclusive, and unlike direct indexing, an out-of-range slice never raises; it just clamps or returns an empty string. A negative step walks backward, which is why s[::-1] is the idiomatic one-line string reversal — there’s no dedicated .reverse() method for strings the way there is for lists.”

Common Mistakes

  • Expecting s[stop] to be included in a slice — stop is always exclusive, so s[2:5] gets indices 2, 3, 4, not 5.
  • Assuming an out-of-range slice raises IndexError like direct indexing does — it silently clamps instead, which can mask a bug where the indices were wrong.
  • Trying s.reverse() — strings have no such method; the idiom is s[::-1].

9.5 String Operators

Concatenation (+)

Joins two strings end-to-end into a new string — the simplest way to build text from pieces, used constantly for messages, file paths, and resource names built from variables.

>>> env, service = "prod", "api"
>>> env + "-" + service + "-server"
'prod-api-server'

Repetition (*)

Repeats a string a given number of times — commonly used for separators, indentation, or simple visual bars without a loop.

>>> "-" * 20
'--------------------'

Membership (in, not in)

Tests whether a substring occurs anywhere inside a string, returning True or False — the simplest way to check for a keyword’s presence without needing find()’s index.

>>> "a" in "Programming", "z" not in "Programming"
(True, True)

Comparison Operators

==, !=, <, >, <=, >= compare strings by their actual character content, not by identity (see 9.1 Introduction to Strings) — this is how exact matches are checked, such as validating a password or a config value.

>>> "password" == "password"
True
>>> "password" != "Password"     # comparison is case-sensitive
True

Lexicographical Comparison

Strings compare character by character using Unicode code point order (dictionary order for ASCII text) — this is what powers sorted() on lists of strings, covered generally in 7.4 Comparison Operators.

>>> "apple" < "banana"
True
>>> "Apple" < "apple"     # uppercase sorts before lowercase (A=65, a=97)
True

Iterating Through Strings

Because a string is an iterable sequence, a for loop visits each character in order — the basis for hand-written parsers, character counters, and validation logic.

>>> for ch in "abc":
...     print(ch, end=" ")
a b c

String Length (len())

len() returns the character count in O(1) time — the length is cached on the object, not recounted each call — used constantly for bounds checks, padding widths, and validation.

>>> len("Programming")
11

Quick Interview Answer

“+ concatenates and * repeats, both producing new strings since strings are immutable. in/not in test substring membership. Comparison operators (==, <, …) compare content, never identity, using Unicode code-point order — the same order sorted() uses on a list of strings. len() is O(1) because the length is cached on the string object rather than recomputed. Iterating a string with for visits one character at a time, since str implements the iterable protocol.”

Common Mistakes

  • Using + to build a large string inside a loop — each += allocates an entirely new string, costing O(n²) overall; see 9.11 Memory and Performance for the join() alternative.
  • Assuming string comparison is case-insensitive — "Password" != "password"; normalize with .lower() first if case shouldn’t matter.
  • Forgetting uppercase letters sort before lowercase ones in code-point order, producing a surprising sort order on mixed-case data.

9.6 String Formatting

Three Ways to Format

flowchart TD A["% operator\nlegacy, printf-style"] --> D["Same output"] B[".format()\nPython 2.7+, readable"] --> D C["f-string\nPython 3.6+, fastest"] --> D

All three produce identical output; f-strings are the recommended default for new code.

% Formatting

The oldest style, borrowed from C’s printf. %s substitutes a string, %d an integer, %f a float. Still seen in older codebases and some logging configs, but largely superseded.

>>> "Server %s is at %s%% capacity" % ("web01", 87)     # %% escapes a literal %
'Server web01 is at 87% capacity'

str.format()

Replaces {} placeholders with arguments, by position or by name. More readable than %-formatting, and still common where the template string itself is loaded from a file or config.

>>> "Server {name} is at {pct}% capacity".format(name="web01", pct=87)
'Server web01 is at 87% capacity'

f-Strings

Embed expressions directly inside {} within the string literal itself, prefixed with f. The fastest and most readable option, and the recommended default for new code since Python 3.6.

>>> name, pct = "web01", 87
>>> f"Server {name} is at {pct}% capacity"
'Server web01 is at 87% capacity'

Alignment and Padding

A format spec’s <, >, and ^ characters left-align, right-align, and center-align a value within a fixed width — the basis for neatly columned console output or reports. Padding fills unused width with a character, spaces by default.

>>> f"{'Alice':<10}|{'30':<5}"     # left-align, width 10 and 5
'Alice     |30   '
>>> f"{7:03d}"                     # zero-pad an integer to width 3
'007'

Number, Currency, and Date Formatting

Format specs insert thousands separators, convert a fraction to a percentage, or apply strftime-style date codes automatically.

>>> f"{1234567:,}"           # thousands separator
'1,234,567'
>>> f"{0.4567:.2%}"          # percentage
'45.67%'
>>> f"${1234.5:,.2f}"        # currency: precision + separator + symbol
'$1,234.50'

>>> import datetime
>>> d = datetime.datetime(2026, 7, 13, 10, 22, 5)
>>> f"{d:%Y-%m-%d %H:%M:%S}"
'2026-07-13 10:22:05'

Logging Format Example

Combining date formatting with a static prefix/suffix is exactly how most log lines are built:

>>> f"[{d:%Y-%m-%d %H:%M:%S}] INFO Service started"
'[2026-07-13 10:22:05] INFO Service started'

Quick Interview Answer

“Python has three formatting styles: %-formatting (legacy, printf-style, error-prone with many arguments), .format() (readable, supports positional and named fields, good when the template itself is loaded as data), and f-strings (fastest, most readable, expressions evaluated inline — the default for new code in 3.6+). Format specs share the same mini-language across all of them for alignment (</>/^), padding, thousands separators, percentages, and strftime-style date codes.”

Common Mistakes

  • Using an f-string on a template string loaded from a config file or user input — f-strings evaluate arbitrary expressions at parse time, so the template itself must be a literal in the source code; use .format() or string.Template when the template is data.
  • Forgetting %% is required to output a literal % in %-style formatting, since a bare % is interpreted as a format specifier.
  • Mixing positional and named .format() arguments incorrectly, or miscounting %-style positional placeholders against the tuple of values, causing a TypeError.

9.7 Common String Methods

Every method below returns a new string (or list/tuple) — none modify the original, consistent with 9.1 Introduction to Strings.

Case Conversion

MethodExampleResult
upper()"Hello".upper()'HELLO'
lower()"Hello".lower()'hello'
capitalize()"hello".capitalize()'Hello'
title()"hello world".title()'Hello World'
swapcase()"Hello".swapcase()'hELLO'
casefold()"STRASSE".casefold()'strasse'

Searching Methods

MethodExampleResult
find()"Hello World".find("World")6
rfind()"Hello World".rfind("o")7
index()"Hello World".index("World")6
count()"banana".count("a")3
startswith()"file.txt".startswith("file")True
endswith()"file.txt".endswith(".txt")True

find() returns -1 if not found; index() raises ValueError. Use find() when “not found” is a normal case, index() when the substring’s absence signals a bug that should surface loudly.

Validation Methods

The is* family answers yes/no questions about a string’s content — used for quick input validation before parsing or type conversion (see 8.1 Introduction to Type Conversion).

MethodExampleResult
isalpha()"abc123".isalpha()False
isdigit()"123".isdigit()True
isalnum()"abc123".isalnum()True
islower() / isupper()"hello".islower()True
isspace()" ".isspace()True
isidentifier()"var_1".isidentifier()True

Modification Methods

MethodExampleResult
replace()"Hello World".replace("World","Python")'Hello Python'
strip()" hi ".strip()'hi'
lstrip() / rstrip()" hi ".lstrip()'hi '
removeprefix()"unit_test.py".removeprefix("unit_")'test.py'
removesuffix()"image.png".removesuffix(".png")'image'

Splitting and Joining Methods

MethodExampleResult
split()"a,b,,c".split(",")['a','b','','c']
rsplit()"a,b,,c".rsplit(",",1)['a,b,','c']
splitlines()"line1\nline2".splitlines()['line1','line2']
partition()"k=v=x".partition("=")('k','=','v=x')
join()",".join(["a","b","c"])'a,b,c'

join() does the reverse of split() — it merges a list of strings into one, inserting the given separator between each piece. It’s the efficient way to build a string from parts (see 9.11 Memory and Performance).

Alignment Methods

MethodExampleResult
center()"hi".center(6,"*")'**hi**'
ljust()"hi".ljust(6,"*")'hi****'
rjust()"hi".rjust(6,"*")'****hi'
zfill()"7".zfill(3)'007'

Encoding and Translation Methods

encode() converts a Unicode str into raw bytes using a given character encoding, usually UTF-8 — required whenever text needs to leave Python as a byte stream, such as writing to a file or sending over a network (see 5.7 Binary Data Types).

>>> "café".encode("utf-8")
b'caf\xc3\xa9'

maketrans() builds a character-mapping table; translate() applies it in one fast pass — useful for bulk character substitution without chaining many replace() calls.

>>> table = str.maketrans("aeiou", "12345")
>>> "hello world".translate(table)
'h2ll4 w4rld'

Quick Interview Answer

“The str method set breaks down into a handful of purpose-based groups: case conversion (upper/lower/title), searching (find/index/count/startswith), validation (the is* family), modification (replace/strip), splitting/joining (split/join — mirror images of each other), alignment (center/ljust/zfill), and encoding (encode/translate). The one recurring theme is that every single one returns a new string rather than mutating the original, since str is immutable.”

Common Mistakes

  • Using index() when “not found” is a normal, expected outcome — it raises ValueError instead of returning a sentinel, so find() is usually the safer default for optional matches.
  • Calling strip() expecting it to remove whitespace from the middle of a string — it only trims from the two ends; use replace(" ", "") or a regex for interior whitespace.
  • Forgetting split() with no arguments splits on any run of whitespace and drops empty strings, while split(",") with an explicit separator does not — "a,,b".split(",") keeps the empty string between the commas.

9.8 String Functions

These are global built-in functions that accept a string as an argument, rather than methods called on the string object itself (len(s) vs. s.upper()).

FunctionExampleResult
len()len("hello")5
max()max("hello")'o'
min()min("hello")'e'
sorted()sorted("dcba")['a','b','c','d']
chr()chr(97)'a'
ord()ord('a')97
ascii()ascii("héllo")"'h\\xe9llo'"
repr()repr("hi\n")"'hi\\n'"
str()str(123)'123'

chr() and ord() are inverses — chr() converts a Unicode code point (an integer) to its character, ord() does the reverse. max()/min() on a string compare characters by their code point, the same ordering used in 9.5 String Operators.

>>> ord("A"), chr(65)
(65, 'A')

repr() produces an unambiguous, developer-facing representation — useful in debugging output because it shows exactly what’s in the string, including otherwise-invisible characters like \n:

>>> print("hi\n")     # str() / print() interprets the escape
hi

>>> print(repr("hi\n"))     # repr() shows it literally
'hi\n'

Quick Interview Answer

“These are functions, not methods — they’re called as len(s) rather than s.len(). len() and sorted()/max()/min() treat a string as an iterable of characters. chr() and ord() convert between a character and its Unicode code point, in either direction. str() converts any object to its readable string form, while repr() produces an unambiguous, code-like representation — the distinction matters most in debugging, where repr() reveals hidden characters like \n that str()/print() would render literally.”

Common Mistakes

  • Calling s.len() instead of len(s) — length is a built-in function, not a string method, unlike most other string operations.
  • Confusing str() and repr() output for debugging — print(value) uses str() and can hide escape sequences or ambiguity that repr(value) would reveal.
  • Assuming max("hello") returns the longest substring — on a string it compares individual characters by code point, returning 'o', not a substring.

9.9 String Algorithms

flowchart TD SA["String Algorithms"] SA --> CH["Check\npalindrome, anagram"] SA --> TR["Transform\nreverse, compress"] SA --> CO["Count\nwords, frequency"] SA --> FI["Find\nduplicates, first non-repeating"]

Each algorithm below reuses tools already covered in this chapter — slicing, sorted(), Counter, and split()/join() — rather than hand-rolled loops wherever Python already provides the idiom.

Reverse String

Extended slicing (see 9.4 String Indexing and Slicing) does it in one line, no loop required.

def reverse_string(s):
    return s[::-1]

>>> reverse_string("hello")
'olleh'

Palindrome Check

A palindrome reads the same forwards and backwards. Checking one means normalizing case and stripping non-alphanumeric characters first, then comparing the string to its own reverse.

def is_palindrome(s):
    s = ''.join(c.lower() for c in s if c.isalnum())
    return s == s[::-1]

>>> is_palindrome("A man a plan a canal Panama")
True

Anagram Check

Two strings are anagrams if they contain exactly the same characters in a different order. Sorting both strings’ characters and comparing is a simple, reliable way to check.

def is_anagram(a, b):
    return sorted(a.lower()) == sorted(b.lower())

>>> is_anagram("listen", "silent")
True

Character Frequency

collections.Counter builds a frequency map of every character in a single pass — useful for statistics, compression, and cipher analysis.

from collections import Counter

>>> dict(Counter("mississippi"))
{'m': 1, 'i': 4, 's': 4, 'p': 2}

Remove Duplicate Characters (Order-Preserving)

dict.fromkeys() removes duplicates while preserving first-seen order — a plain set() would not, since sets don’t preserve insertion order.

def remove_duplicates(s):
    return "".join(dict.fromkeys(s))     # dict preserves first-seen order

>>> remove_duplicates("programming")
'progamin'

Word and Character Statistics

split() with no arguments splits on any run of whitespace, so counting the resulting list’s length is the simplest way to count words:

>>> len("the quick brown fox".split())     # word count
4
>>> sum(1 for c in "DevOps".lower() if c in "aeiou")     # vowel count
2

Reversing the order of words (different from reversing characters) means splitting into words, reversing the list, and rejoining:

>>> " ".join("the sky is blue".split()[::-1])
'blue is sky the'

Longest / Shortest Word

max()/min() with key=len finds the “biggest” or “smallest” item by any custom measure, not just length — a pattern that generalizes well beyond strings.

>>> words = "Kubernetes simplifies container orchestration".split()
>>> max(words, key=len), min(words, key=len)
('orchestration', 'simplifies')

First Non-Repeating Character

A classic interview warm-up: find the first character that appears exactly once.

def first_non_repeating(s):
    for c in s:
        if s.count(c) == 1:
            return c
    return None

>>> first_non_repeating("swiss")
'w'

String Compression (Run-Length Encoding)

Collapse runs of repeated characters into char+count pairs, falling back to the original if compression doesn’t actually help.

def compress(s):
    result, i = [], 0
    while i < len(s):
        j = i
        while j < len(s) and s[j] == s[i]:
            j += 1
        result.append(s[i] + str(j - i))
        i = j
    compressed = "".join(result)
    return compressed if len(compressed) < len(s) else s

>>> compress("aaabbbcca")
'a3b3c2a1'

Quick Interview Answer

“Most string algorithm questions reduce to combining a small set of tools rather than hand-rolled character loops: s[::-1] for reversal, comparing a cleaned string to its own reverse for palindromes, sorted() for anagrams, Counter for frequency counting, and dict.fromkeys() for order-preserving deduplication. s.count(c) inside a loop finds the first non-repeating character in O(n²) in the naive form; a Counter pass first makes it O(n) — worth mentioning if performance comes up. Run-length compression is the one genuinely algorithmic pattern here: scan for runs, encode char+count, and fall back to the original if compression doesn’t actually shrink it.”

Common Mistakes

  • Using a plain set() to deduplicate characters when order matters — sets don’t preserve insertion order; dict.fromkeys(s) does.
  • Writing first_non_repeating with s.count(c) inside the loop for a very long string without realizing it’s O(n²) — a Counter pre-pass makes it O(n).
  • Forgetting a palindrome check needs to strip punctuation and normalize case first — "A man, a plan" naively fails a raw s == s[::-1] check.

9.10 Regular Expressions with Strings

What Is Regex?

A regular expression is a mini-language for describing text patterns — used for validation, extraction, and bulk find/replace. Python’s built-in re module implements it, and it’s the production-grade alternative to the naive character-by-character search shown in 9.9 String Algorithms for anything beyond simple literal matching.

flowchart LR G["(?P<year>\\d{4})-(?P<month>\\d{2})-(?P<day>\\d{2})"] G --> Y["year group\n\\d{4}"] G --> M["month group\n\\d{2}"] G --> D["day group\n\\d{2}"]

A named-group pattern for 2026-07-13: match.group("year") → "2026".

re.match() vs re.search()

match() only succeeds if the pattern matches starting at position 0; search() scans the whole string for the first match anywhere. Use match() to validate that an entire string starts a certain way, search() to find something buried inside larger text.

>>> import re
>>> re.match(r"\d+", "123abc")        # anchored at the START of the string
<re.Match object; span=(0, 3), match='123'>
>>> re.search(r"\d+", "abc123def")    # finds the FIRST match anywhere
<re.Match object; span=(3, 6), match='123'>

re.findall() / re.finditer()

findall() collects every match into a list of strings; finditer() gives the same matches lazily as Match objects, useful when each match’s position is also needed.

>>> re.findall(r"\d+", "a1 b22 c333")
['1', '22', '333']

re.sub() / re.split()

sub() replaces every match with a given string — the regex equivalent of str.replace() but pattern-based. split() breaks a string apart wherever the pattern matches, useful when the separator itself is irregular (variable amounts of whitespace, for example).

>>> re.sub(r"\s+", " ", "too    many   spaces")
'too many spaces'
>>> re.split(r",\s*", "a, b,c,  d")
['a', 'b', 'c', 'd']

Pattern Compilation

Compiling a pattern once and reusing it for repeated matching (e.g. inside a loop over thousands of log lines) avoids re-parsing the pattern every call.

>>> pattern = re.compile(r"\d+")
>>> pattern.findall("a1 b22")
['1', '22']

Groups and Named Groups

Parentheses in a pattern capture matched sub-text for later use; groups() returns all captures as a tuple in order. (?P<name>...) attaches a label so a group can be retrieved by name instead of position — far more readable with many groups, and less fragile to reorder.

>>> m = re.match(r"(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})", "2026-07-13")
>>> m.group("year"), m.group("month"), m.group("day")
('2026', '07', '13')

Lookahead and Lookbehind

A lookahead (?=...) matches a position only if followed by a given pattern, without including that pattern in the match — useful for a value only when it has a specific unit or suffix nearby. A lookbehind (?<=...) is the mirror image.

>>> re.findall(r"\d+(?=px)", "10px 20em 30px")     # digits followed by 'px'
['10', '30']
>>> re.findall(r"(?<=\$)\d+", "$50 and 30")        # digits preceded by '$'
['50']

Quick Interview Answer

“match() anchors at position 0; search() finds a match anywhere. findall() returns matched strings; finditer() returns lazy Match objects with position info. sub() and split() are the pattern-based equivalents of str.replace() and str.split(). Capture groups pull structured sub-fields out of a match — named groups ((?P<name>...)) make multi-group patterns readable instead of positional and fragile. Lookahead/lookbehind assert that a pattern exists nearby without consuming it into the match itself. Compiling a pattern once with re.compile() avoids re-parsing it on every call inside a hot loop.”

Common Mistakes

  • Using match() when search() is actually needed — match() silently returns None for anything not anchored at position 0, which is easy to misdiagnose as “the pattern is wrong.”
  • Forgetting parentheses in a pattern create a capturing group even when capture isn’t the intent — use (?:...) for a non-capturing group if grouping is only needed for precedence.
  • Re-compiling the same pattern inside a loop instead of compiling it once outside — wasteful when processing many lines, as in log parsing (see 9.12 Strings in DevOps and AWS).

9.11 Memory and Performance

Why Immutability Matters Here

Immutability makes strings safe to share across functions, threads, and dict/set keys without defensive copying, and it’s what allows CPython to intern and cache strings safely (see 9.1 Introduction to Strings). It’s also the direct cause of the single most important string performance rule below.

String Memory Optimization

Since PEP 393, CPython stores each string at 1, 2, or 4 bytes per character depending on its widest code point — a pure-ASCII log line costs far less memory than one containing emoji or CJK text.

Efficient Concatenation: join() as StringBuilder

Because strings are immutable, each += inside a loop allocates a brand-new string and copies everything seen so far into it — costing O(n²) overall for n pieces. join() instead collects all the pieces first and allocates the final buffer exactly once, costing O(n).

# Inefficient: allocates a new string object every iteration -- O(n^2)
result = ""
for word in word_list:
    result += word + " "

# Efficient: build a list, join once -- O(n)
result = " ".join(word_list)

When to Use f-Strings

  • Default choice for readability and speed in Python 3.6+.
  • Prefer .format() when the template string itself is data (loaded from a config file) — see 9.6 String Formatting.
  • Prefer string.Template when the template comes from an untrusted source, since f-strings evaluate arbitrary expressions.

Time Complexity of String Operations

Knowing the Big-O cost of each operation is what separates code that works on a small test from code that stays fast on a million-line log file:

OperationTime ComplexityNotes
Indexing s[i]O(1)Direct memory offset
Slicing s[a:b]O(k)k = length of the slice
Concatenation s1 + s2O(n)n = combined length; O(n²) if repeated in a loop
len(s)O(1)Length is cached on the object
in / not inO(n)Linear scan
str.join(list)O(n)n = total combined length
str.replace() / .split()O(n)Single pass
hash(s)O(n) first time, O(1) afterHash is cached on the object

Quick Interview Answer

“The single most important performance rule for strings: building one incrementally with += inside a loop is O(n²), because every += allocates a brand-new string and copies everything accumulated so far. ''.join(pieces) is O(n) instead, since it computes the total length once and allocates exactly one buffer. This follows directly from immutability — there’s no in-place append the way a list has. Beyond that, indexing and len() are O(1) since strings cache their length, while in, replace(), and split() are all O(n) linear scans.”

Common Mistakes

  • Building a large string with += in a loop instead of accumulating pieces in a list and calling "".join(...) once at the end.
  • Assuming len(s) re-scans the string each call — it’s O(1) because the length is cached on the object at creation time.
  • Ignoring memory cost differences between ASCII and wide-Unicode strings when processing huge volumes of text — a string full of emoji or CJK characters can use 2-4x the memory of an equivalent-length ASCII string.

9.12 Strings in DevOps and AWS

Almost everything touched in cloud and infrastructure automation is a string: log lines, Amazon Resource Names (ARNs), IAM policy documents, kubectl output, and CI console text. Reliable parsing of these formats is a core DevOps skill, and it all follows the same shape.

flowchart LR A["Raw Text\nlog line / ARN / JSON blob"] --> B["Split or Regex Match\n.split() / re.match() / json.loads()"] B --> C["Structured Fields\ndict / tuple / named groups"] C --> D["Analysis / Action\nfilter, count, alert, report"]

The standard text-processing pipeline behind almost every log-parsing or infrastructure-automation script.

Reading and Parsing Log Files

Iterating a file object line-by-line is memory-efficient (it doesn’t load the whole file at once) and is the standard entry point for any log-processing script.

with open("app.log") as f:
    for line in f:
        if "ERROR" in line:
            print(line.strip())

Apache’s “combined” log format is a fixed layout — a single regex with capture groups (see 9.10 Regular Expressions with Strings) extracts every field in one pass; Nginx’s default combined format is close enough that the same pattern style applies:

>>> import re
>>> log = '127.0.0.1 - - [13/Jul/2026:10:22:05 +0000] "GET /index.html HTTP/1.1" 200 1024'
>>> pattern = r'(\S+) \S+ \S+ \[(.*?)\] "(\S+) (\S+) (\S+)" (\d+) (\d+)'
>>> re.match(pattern, log).groups()
('127.0.0.1', '13/Jul/2026:10:22:05 +0000', 'GET', '/index.html', 'HTTP/1.1', '200', '1024')

Linux syslog-style lines (timestamp, hostname, process[pid]: message) follow a similarly predictable shape:

>>> syslog = "Jul 13 10:22:05 web01 sshd[1234]: Accepted publickey for admin"
>>> re.match(r"(\w+ +\d+ [\d:]+) (\S+) (\w+)\[(\d+)\]: (.+)", syslog).groups()
('Jul 13 10:22:05', 'web01', 'sshd', '1234', 'Accepted publickey for admin')

Structured Config Formats

JSON, YAML, and CSV are the standard data-interchange and config formats in DevOps — always prefer the dedicated parser (json, yaml/PyYAML, csv) over hand-rolled string splitting, since they correctly handle quoting, escaping, and nested structure that manual parsing gets wrong.

>>> import json
>>> json.loads('{"a":1,"b":[1,2,3]}')
{'a': 1, 'b': [1, 2, 3]}

>>> import yaml     # requires: pip install pyyaml
>>> yaml.safe_load("name: web01\nport: 8080")
{'name': 'web01', 'port': 8080}

Amazon Resource Names (ARNs)

An ARN has the fixed shape arn:partition:service:region:account-id:resource. Because the resource portion can itself contain colons, split with a maxsplit of 5:

>>> arn = "arn:aws:lambda:us-east-1:123456789012:function:my-function"
>>> arn.split(":", 5)
['arn', 'aws', 'lambda', 'us-east-1', '123456789012', 'function:my-function']

Amazon EC2 and Amazon S3 Identifiers

Amazon EC2 instance IDs follow a fixed i-<hex digits> pattern, easy to pull out of free-form console text with a targeted regex. Amazon S3 virtual-hosted-style URLs embed the bucket name as a subdomain:

>>> text = "Instance i-0abcd1234ef567890 launched in us-east-1a"
>>> re.search(r"i-[0-9a-f]{8,17}", text).group()
'i-0abcd1234ef567890'

>>> url = "https://my-bucket.s3.amazonaws.com/path/to/object.txt"
>>> re.match(r"https://([^.]+)\.s3", url).group(1)
'my-bucket'

AWS CloudWatch Logs

Lambda and other AWS service log lines embed a RequestId and timing data inline — extracting them lets you correlate and measure invocation performance across thousands of log entries. Full working scripts for this pattern live in the Python Scripting section, such as the AWS CloudTrail Root-Account Monitor.

>>> cw_log = "2026-07-13T10:22:05.123Z [INFO] RequestId: abc-123 Duration: 45.67 ms"
>>> re.search(r"RequestId: (\S+) Duration: ([\d.]+) ms", cw_log).groups()
('abc-123', '45.67')

Kubernetes Pod Names

A Deployment-managed pod name has the shape <deployment>-<replicaset-hash>-<pod-suffix> — the two hashes are generated by the ReplicaSet and the kubelet, not the Deployment directly.

>>> pod = "nginx-deployment-66b6c48dd5-x7z2p"
>>> pod.rsplit("-", 2)
['nginx-deployment', '66b6c48dd5', 'x7z2p']

Terraform Output and Git Metadata

terraform apply prints resource IDs inline as it provisions infrastructure — capturing them lets a script feed newly created resource IDs into the next pipeline step. Team branch-naming and commit-message conventions can similarly be parsed to auto-link commits to issue trackers.

>>> tf_out = 'aws_instance.web: Creation complete after 45s [id=i-0abcd1234ef567890]'
>>> re.search(r"\[id=([\w-]+)\]", tf_out).group(1)
'i-0abcd1234ef567890'

>>> commit = "feat(auth): add OAuth2 login support"     # Conventional Commits format
>>> re.match(r"(\w+)\(([^)]+)\): (.+)", commit).groups()
('feat', 'auth', 'add OAuth2 login support')

Quick Interview Answer

“In DevOps and AWS automation, strings are the interface everywhere — logs, ARNs, kubectl output, and CI console text are all just text a script must parse reliably. The pattern is always the same pipeline: raw text in, a split() or regex match extracts structured fields, then those fields drive analysis or action. For fixed-delimiter formats like ARNs, split() with a maxsplit is enough. For log lines with a predictable-but-not-purely-delimited shape, a regex with named capture groups is the standard tool. For genuinely structured data — JSON, YAML, CSV — always reach for the dedicated parser instead of hand-rolled string splitting, since edge cases like quoting and escaping are easy to get subtly wrong by hand.”

Common Mistakes

  • Splitting an ARN with a plain .split(":") and no maxsplit — the resource portion can itself contain colons, silently breaking apart a field that should have stayed intact.
  • Hand-parsing JSON or YAML with string methods instead of json.loads()/yaml.safe_load() — quoting, escaping, and nesting are easy to get subtly wrong by hand.
  • Treating a Kubernetes pod name’s suffix as meaningful or stable — it’s a randomly generated hash from the ReplicaSet and kubelet, not something to parse for business logic.

9.13 Best Practices

Build Large Strings with join(), Not +=

Accumulating a string with += inside a loop is O(n²); collecting pieces in a list and calling "".join(...) once is O(n). See 9.11 Memory and Performance for the full explanation.

result = " ".join(word_list)     # preferred over looped += 

Default to f-Strings, but Know the Exceptions

f-strings are the fastest and most readable choice for new code — but only when the template itself is a literal in the source. If the template is loaded from a config file or comes from an untrusted source, use .format() or string.Template instead (see 9.6 String Formatting).

Compare Content with ==, Never Identity with is

String interning is a CPython implementation detail, not a language guarantee — two equal strings built at runtime are not guaranteed to be the same object. Always compare string content with ==.

>>> a = "hello"; b = "".join(["h", "e", "l", "l", "o"])
>>> a == b, a is b
(True, False)     # equal content, but NOT the same object

Parse Structured Data with the Standard Library, Not eval()

Use json.loads(), ast.literal_eval(), or yaml.safe_load() to parse data — never eval(), which executes arbitrary code and is a serious security risk on any input that isn’t fully trusted.

Specify Encoding Explicitly

Always pass an explicit encoding="utf-8" (or whatever is correct) when reading or writing files, rather than relying on the platform default — the default varies by OS and can silently corrupt non-ASCII text.

Quick Interview Answer

“Five habits cover most of what matters day to day: build large strings with ''.join(...) instead of += in a loop; default to f-strings for new code, but fall back to .format()/string.Template when the template itself is data or untrusted; always compare strings with ==, never is, since interning isn’t a language guarantee; parse structured text with json.loads()/ast.literal_eval() rather than eval(), which is a security risk on untrusted input; and specify file encoding explicitly instead of relying on a platform default that can silently corrupt non-ASCII text.”

Common Mistakes

  • Reaching for eval() to parse a dict-like or list-like string instead of ast.literal_eval() or json.loads() — eval() executes arbitrary code, which is unsafe on anything not fully trusted.
  • Opening a file with open(path) and no explicit encoding=, then hitting a UnicodeDecodeError or silent mojibake on a different OS or locale.
  • Relying on is for string comparison because it happened to work in a quick REPL test — interning behavior differs between short literals and runtime-built strings.

9.14 Common Mistakes

Assuming a Method Mutates in Place

Every string method returns a new string — the original is never changed, since str is immutable (see 9.1 Introduction to Strings).

>>> s = "hello"
>>> s.upper()          # returns a new string
>>> s                  # s itself is UNCHANGED
'hello'
>>> s = s.upper()      # must reassign to actually keep the result

Confusing find() and index()

find() returns -1 when nothing matches; index() raises ValueError. Using the wrong one for the situation either silently produces a bogus -1 result or crashes unexpectedly.

>>> "hello".find("z")      # no exception -- easy to silently misuse
-1
>>> "hello".index("z")     # raises instead
Traceback (most recent call last):
ValueError: substring not found

Off-by-One Slicing Errors

Forgetting that stop in s[start:stop] is exclusive is one of the most common slicing bugs — s[0:5] gets 5 characters (indices 0–4), not 6.

O(n²) Concatenation in a Loop

Building a large string with += inside a loop looks harmless on small inputs but degrades badly at scale — see 9.11 Memory and Performance for why, and join() for the fix.

# Looks fine on 10 items, degrades badly on 100,000
result = ""
for line in lines:
    result += line

Comparing Strings with is

Interning is a CPython implementation detail, not a guarantee — comparing with is can pass in a quick test and fail unpredictably on runtime-built strings.

>>> a = "hello"
>>> b = "".join(["h", "e", "l", "l", "o"])
>>> a == b       # correct: compares content
True
>>> a is b       # WRONG tool: compares identity, not guaranteed True
False

Quick Interview Answer

“The recurring string bugs share one root cause or another: forgetting immutability means a method’s return value must be captured (s = s.upper(), not just s.upper()); mixing up find()’s silent -1 with index()’s ValueError produces either a bogus result or an unexpected crash; forgetting slicing’s stop is exclusive causes off-by-one errors; building a large string with += in a loop is O(n²) and only shows up as a real problem at scale; and comparing with is instead of == relies on CPython’s interning, which isn’t a language guarantee.”

Common Mistakes

  • Writing s.strip() on a line and expecting s itself to be stripped afterward — the call must be assigned back: s = s.strip().
  • Using find()’s -1 return value directly as an index without checking for it first, silently slicing from the end of the string instead of failing loudly.
  • Reaching for += string-building inside a loop that turns out to run over a large or unbounded number of iterations (a log file, an API pagination loop) — the O(n²) cost compounds fast.

9.15 Interview Questions

Conceptual Questions

  • Why are Python strings immutable, and what does that mean for methods like .replace()?
  • What’s the difference between find() and index()?
  • Why is s[::-1] the idiomatic way to reverse a string when there’s no .reverse() method?
  • Why is building a string with += in a loop worse than using .join()?
  • What’s the difference between str.format() and an f-string?

Scenario-Based Questions

  • A script processes a multi-gigabyte log file and needs to build a filtered output string line by line — how should it be built to avoid a performance cliff?
  • An API response sometimes has extra whitespace or inconsistent casing in a status field — how would you normalize it reliably before comparing it?
  • A regex pattern used inside a hot loop over thousands of log lines is running slowly — what’s the first optimization to check?

Coding Questions

  • Check whether a given string is a palindrome, ignoring case and punctuation.
  • Check whether two strings are anagrams of each other.
  • Find the longest common prefix shared by a list of strings.
  • Find the first non-repeating character in a string.
  • Implement basic run-length string compression.
def longest_common_prefix(strs):
    if not strs:
        return ""
    prefix = strs[0]
    for s in strs[1:]:
        while not s.startswith(prefix):
            prefix = prefix[:-1]
    return prefix

>>> longest_common_prefix(["flower", "flow", "flight"])
'fl'

Quick Interview Answer

“String interview questions cluster around three things: confirming immutability is understood (every method returns a new object; .replace() on its own does nothing observable), whether the standard idioms are known instead of hand-rolled loops (s[::-1] for reversal, sorted(a) == sorted(b) for anagrams, ''.join(...) over += for building large strings), and classic coding problems — palindrome, anagram, longest common prefix, first non-repeating character — that test comfort combining slicing, Counter, and sorted() rather than writing everything from scratch.”

Common Mistakes

  • Answering “strings are immutable so they can’t be modified at all” without the follow-up: every method still returns a new string, so s.strip() is very much useful — just not in place.
  • Solving the longest-common-prefix question with nested loops comparing every character position across every string, instead of the simpler shrink-a-candidate-prefix approach.
  • Reaching for a manual character-count loop for the anagram check instead of the one-line sorted(a) == sorted(b) idiom.

9.16 Hands-on Exercises

Log Level Analyzer

Build a function that scans a batch of log lines and tallies how many fall into each severity level — an at-a-glance health summary without opening a log viewer.

import re
from collections import Counter

def analyze_log_levels(lines):
    levels = [re.search(r"\b(INFO|WARNING|ERROR)\b", l).group(1) for l in lines]
    return dict(Counter(levels))

>>> analyze_log_levels([
...     "2026-07-13 10:00:01 INFO Service started",
...     "2026-07-13 10:00:05 ERROR Connection refused",
...     "2026-07-13 10:00:07 WARNING High memory usage",
... ])
{'INFO': 1, 'ERROR': 1, 'WARNING': 1}

IP Address Extractor

Pull every IPv4 address out of free-form text — the first step in building a blocklist, an allowlist audit, or a geo-lookup report from security logs.

>>> text = "Connections from 192.168.1.10 and 10.0.0.5 were blocked; 8.8.8.8 allowed"
>>> re.findall(r"\b(?:\d{1,3}\.){3}\d{1,3}\b", text)
['192.168.1.10', '10.0.0.5', '8.8.8.8']

Configuration File Parser

Read a simple key=value config format, with comment lines starting with # — a recurring task for any tool that needs its own settings file without pulling in a full config-parsing library.

def parse_config(text):
    result = {}
    for line in text.splitlines():
        line = line.strip()
        if not line or line.startswith("#"):
            continue
        key, _, value = line.partition("=")
        result[key] = value
    return result

>>> parse_config("# config\nhost=localhost\nport=8080\n\ndebug=true")
{'host': 'localhost', 'port': '8080', 'debug': 'true'}

Password Strength Checker

Enforce a minimum length plus a mix of character classes — a standard first line of defense against weak passwords during account signup.

def is_strong_password(pw):
    if len(pw) < 8:
        return False
    has_upper = any(c.isupper() for c in pw)
    has_lower = any(c.islower() for c in pw)
    has_digit = any(c.isdigit() for c in pw)
    has_special = any(not c.isalnum() for c in pw)
    return all([has_upper, has_lower, has_digit, has_special])

>>> is_strong_password("Weak1")
False
>>> is_strong_password("Str0ng!Pass")
True

Mini Projects

  • DevOps health report generator — turn raw HTTP status-code counts into a single human-readable error-rate line ("1000 requests, 2.0% error rate"), reusing Counter from earlier in this chapter.
  • CI/CD build summarizer — parse a Jenkins-style console line ("Build #42 SUCCESS in 3m 15s") into a structured {"build", "status", "duration"} result, ready to feed a Slack notification.
  • Kubernetes log analyzer — combine the pod-name parser from 9.12 Strings in DevOps and AWS with the log-level counter above to produce a per-deployment error-rate summary.

Quick Interview Answer

“These exercises tie the chapter’s tools together: the log analyzer combines regex extraction with Counter; the IP extractor is a single well-chosen regex; the config parser leans on partition() instead of a fragile manual split; and the password checker chains several is* validation methods with any(). The common thread is picking the right existing tool — regex, Counter, partition, is* — over hand-rolled character-by-character logic.”

Common Mistakes

  • Using split("=") instead of partition("=") in the config parser — split() breaks on a value that itself contains an = character, while partition() splits only on the first occurrence.
  • Forgetting the password checker’s has_special check needs not c.isalnum(), not a hardcoded set of symbols, so it correctly accepts any real special character.
  • Letting one malformed log line crash the whole analyzer instead of skipping or flagging lines that don’t match the expected pattern (re.search(...) returning None).

Strings: Chapter Practice

Mini lab

Normalize " WEB-01 " for a case-insensitive lookup.

Expected output: web-01

HintUse the chapter examples to choose the operation, then run your version before opening the solution.
Show solution
name = "  WEB-01  "
print(name.strip().lower())

Knowledge check

Does strip() modify the original string?

Check your answerNo. String methods return new strings.

Finish the chapter

  • Explain each operation without reading the lesson.
  • Change the sample input and predict the output before running it.
  • Revisit the chapter if the result differs from your prediction.

Back to this chapter · Choose the next chapter

10.1 Introduction to Lists

Python Lists

What Is a List?

What Is It?

An ordered, mutable collection that can hold any mix of values, written with square brackets and comma-separated.

>>> servers = ["web01", "web02", "db01"]
>>> servers
['web01', 'web02', 'db01']

Why Does It Matter?

A list is Python’s general-purpose, resizable “array” — the default choice whenever an ordered group of items might grow, shrink, or be reordered, unlike the fixed-content tuple (see 5.4 Sequence Data Types).

Characteristics

  • Ordered — items keep the position they were inserted in.
  • Mutable — can be changed in place after creation (see 10.11 Copying and Mutability).
  • Allows duplicates — the same value can appear more than once.
  • Heterogeneous — can mix types freely (int, str, even other lists).

Why Lists?

Lists are the right tool whenever order matters and the collection’s size isn’t fixed in advance — the majority of everyday collections (a queue of tasks, a batch of records, a sequence of pipeline steps) fit this shape.

Real-World Applications

  • A shopping cart’s items
  • Rows returned from a database query
  • A sequence of steps in a pipeline
  • A batch of files to process

Lists in DevOps

Lists are everywhere in infrastructure scripting: server inventories, Amazon EC2 instance IDs, container names, log lines — 10.15 Lists in DevOps is dedicated entirely to these patterns. (Time complexity for append() is covered fully in 10.13 Performance.)

>>> allowed_regions = ["us-east-1", "us-west-2", "eu-west-1"]

Internal Representation

flowchart LR L["variable l"] --> A["contiguous array of POINTERS"] A --> P1["→ 1"] A --> P2["→ 2"] A --> P3["→ 3"] A --> S1["(spare)"] A --> S2["(spare)"]

Each slot is a reference to an object elsewhere on the heap, with spare capacity pre-allocated for fast appends.

CPython allocates a contiguous block of memory holding pointers to each element — not the elements’ actual data, which lives separately on the heap (see 6.3 Memory Management: Stack vs Heap). Every slot is a same-sized reference, exactly like a variable, which is why a single list can freely hold mixed types.

>>> import sys
>>> sys.getsizeof([])       # empty list overhead
56
>>> sys.getsizeof([1])      # one element added
88

Dynamic Memory and Capacity

Unlike a fixed-size C array, a Python list resizes itself automatically — no capacity is ever declared up front. When it needs to grow, CPython over-allocates a few extra, unused slots so the next several append() calls don’t need another resize. This amortizes the cost of growing: append() is O(1) on average even though an occasional resize is O(n) — see 10.13 Performance.

>>> l = [1, 2, 3]
>>> id(l)
140632280611584
>>> l.append(4)     # fits in spare capacity -- no reallocation needed
>>> id(l)            # SAME id -- same underlying array object
140632280611584

Quick Interview Answer

“A list is an ordered, mutable, heterogeneous collection — Python’s general-purpose resizable array. Internally, CPython stores it as a contiguous array of pointers to the actual element objects, which live separately on the heap; that’s why a single list can mix types freely, since every slot is just a same-sized reference regardless of what it points to. To make growth cheap, CPython over-allocates spare capacity whenever it resizes, which is exactly what makes append() O(1) amortized instead of paying a full reallocation cost on every single call.”

Common Mistakes

  • Assuming a list stores its elements’ actual values contiguously, like a C array — it stores contiguous pointers; the values themselves live elsewhere on the heap.
  • Reaching for a list when the collection is genuinely fixed and never changes — a tuple is more appropriate and slightly more memory-efficient (see 5.4 Sequence Data Types).
  • Expecting append() to be uniformly O(1) with zero variance — it’s O(1) amortized; the occasional resize underneath is O(n), it’s just rare enough not to matter in aggregate.

10.2 Creating Lists

Empty Lists

The starting point for building up a collection incrementally, e.g. inside a loop.

>>> items = []
>>> items
[]

List Literals

The most common way to create a list with known contents up front — square brackets with comma-separated values.

>>> nums = [1, 2, 3]

list()

The list() constructor converts any iterable — string, tuple, range, set — into a list, covered in depth in 8.6 Collection Type Conversion.

>>> list((1, 2, 3))
[1, 2, 3]

Lists from Strings

Passing a string to list() splits it into its individual characters — a different result than str.split(), which splits by word or delimiter (see 9.7 Common String Methods).

>>> list("abc")
['a', 'b', 'c']

range()

range() alone is lazy — wrapping it in list() actually materializes the numbers as a real list (see 5.4 Sequence Data Types).

>>> list(range(5))
[0, 1, 2, 3, 4]

Nested Lists

A list can contain other lists as elements — the basis for representing grids, matrices, and tables; see 10.10 Nested Lists and Matrices.

>>> grid = [[1, 2], [3, 4]]

Mixed Data Types

Unlike arrays in many other languages, a single Python list can freely mix types.

>>> mixed = [1, "two", 3.0, True]

Quick Interview Answer

“A list literal ([1, 2, 3]) is the common case for known contents. list() converts any iterable — a string becomes its individual characters, not words, which trips people up if str.split() was actually intended. range() is lazy on its own; list(range(...)) is what actually materializes the numbers. Lists can nest freely and mix types, since every slot is just a same-sized reference regardless of what it points to.”

Common Mistakes

  • Calling list("hello world") expecting a list of words — it produces a list of individual characters, including the space; str.split() is the word-splitting tool.
  • Materializing a large range() into a list() unnecessarily, losing its lazy memory efficiency for no reason.
  • Writing [[0] * 3] * 3 to build a 3×3 grid of independent rows — * on a list repeats references to the same inner list, so mutating one row mutates all of them (see 10.10 Nested Lists and Matrices for the correct pattern).

10.3 List Indexing and Slicing

Indexing

flowchart LR A["10\n0 / -5"] --- B["20\n1 / -4"] --- C["30\n2 / -3"] --- D["40\n3 / -2"] --- E["50\n4 / -1"]

Positive indices count from the left; negative indices count from the right.

Positive indices count from the left, starting at 0. Negative indices count from the right, starting at -1 — reaching the end without needing len(l) - 1.

>>> l = [10, 20, 30, 40, 50]
>>> l[0], l[2]
(10, 30)
>>> l[-1]
50

Nested Indexing

Chain indices to reach into nested lists — the first index selects the inner list, the second selects an element within it.

>>> nested = [[1, 2], [3, 4]]
>>> nested[0][1]
2

IndexError

Accessing an index outside the list’s actual range raises IndexError immediately — unlike slicing below, which never does.

>>> l[10]
Traceback (most recent call last):
IndexError: list index out of range

Slicing

l[start:stop:step] extracts a sub-list — the exact same syntax as string slicing (see 9.4 String Indexing and Slicing), including the same “never raises” behavior for out-of-range bounds.

>>> l = [10, 20, 30, 40, 50]
>>> l[1:3]      # start:stop
[20, 30]
>>> l[:2]       # start defaults to 0
[10, 20]
>>> l[2:]       # stop defaults to the end
[30, 40, 50]
>>> l[::2]      # every 2nd element
[10, 30, 50]

Reverse

A negative step of -1 reverses the list, mirroring the string idiom exactly.

>>> l[::-1]
[50, 40, 30, 20, 10]

Copy

l[:] creates a new list object — unlike the equivalent string slice, which CPython can safely reuse because strings are immutable (see 9.4 String Indexing and Slicing). Lists cannot be shared this way, so a genuine copy is always made.

>>> copy_l = l[:]
>>> copy_l is l     # a DIFFERENT object, not the same one
False

l[:] is a shallow copy — fine for a flat list, but nested lists inside are still shared. See 10.11 Copying and Mutability for the full shallow-vs-deep comparison.

Slicing as an Assignment Target

Unlike a string slice, a list slice can be assigned to — replacing a whole range of elements at once, covered in 10.4 Modifying Lists.

Quick Interview Answer

“Indexing and slicing on a list use the exact same l[start:stop:step] syntax as strings — the difference is what they hand back and what they allow. Direct indexing out of range raises IndexError; slicing never does, it just clamps. l[::-1] reverses, same as strings. The one list-specific difference: l[:] genuinely copies the list, because lists are mutable and can’t be safely shared the way an immutable string slice can — and unlike a string, a list slice can also be used as an assignment target to replace a whole range in place.”

Common Mistakes

  • Assuming l[:] and the equivalent string slice behave identically under the hood — for strings CPython can return the same object; for lists it must always allocate a new one, since the result could otherwise be mutated unexpectedly.
  • Forgetting l[:] is only a shallow copy — a nested list inside is still the same shared object in both the copy and the original.
  • Expecting an out-of-range slice to raise, the way direct indexing does — l[100:200] on a short list just returns [], no error.

10.4 Modifying Lists

All of these mutate the list in place — the list’s identity (id()) stays the same throughout; see 10.11 Copying and Mutability.

Update

Assign directly to an index to replace that single element.

>>> l = [1, 2, 3]
>>> l[0] = 99
>>> l
[99, 2, 3]

Append

Adds one item to the end — the most common way to grow a list.

>>> l.append(4)
>>> l
[99, 2, 3, 4]

Extend

Adds each item from another iterable individually — different from append(), which would add the whole iterable as one nested element.

>>> l.extend([5, 6])
>>> l
[99, 2, 3, 4, 5, 6]

Insert

Adds an item at a specific position, shifting everything after it one slot to the right — see 10.13 Performance for why this is more expensive than append().

>>> l.insert(1, 100)
>>> l
[99, 100, 2, 3, 4, 5, 6]

Replace (Slice Assignment)

Assigning to a slice replaces a whole range of elements at once, and the replacement can even be a different length than the original range.

>>> l[1:3] = [7, 8, 9]
>>> l
[99, 7, 8, 9, 3, 4, 5, 6]

Quick Interview Answer

“Every modification method here mutates the list in place rather than returning a new one — the list’s id() never changes. l[i] = x replaces a single element. append(x) adds exactly one item to the end; extend(iterable) unpacks and adds each item individually — a very common point of confusion, since append([1, 2]) nests a whole list as one element instead of adding 1 and 2 separately. insert(i, x) shifts everything after position i to make room. Slice assignment is the most powerful of the group — it can replace, grow, or shrink a range in one statement, since the replacement doesn’t need to match the original range’s length.”

Common Mistakes

  • Calling l.append([1, 2]) when l.extend([1, 2]) was intended — append nests the whole list as a single element, producing [..., [1, 2]] instead of adding 1 and 2 individually.
  • Using l.insert(0, x) repeatedly in a loop instead of building the list in order and calling append() — each insert(0, ...) is O(n), turning the loop into O(n²).
  • Forgetting slice assignment’s replacement length doesn’t need to match the slice’s length — l[1:3] = [7, 8, 9] replaces two elements with three, growing the list.

10.5 Removing Elements

remove()

Removes the first occurrence of a given value (not a position) — raises ValueError if the value isn’t present.

>>> l = [1, 2, 3, 2, 1]
>>> l.remove(2)     # removes the first '2' only
>>> l
[1, 3, 2, 1]

pop()

Removes and returns the element at a given index (default: the last one) — the only removal method that gives the removed value back.

>>> popped = l.pop()     # removes and returns the last element
>>> popped, l
(1, [1, 3, 2])
>>> l.pop(0)              # remove and return by specific index
1

clear()

Empties the list completely, in place — the list object still exists (same id()), just with zero elements.

>>> l2 = [1, 2, 3]
>>> l2.clear()
>>> l2
[]

del

A statement (not a method) that removes an item by index — or, combined with a slice, an entire range at once.

>>> l3 = [1, 2, 3, 4, 5]
>>> del l3[0]
>>> l3
[2, 3, 4, 5]
>>> del l3[1:3]     # delete by slice
>>> l3
[2, 5]

Quick Interview Answer

“Four tools cover removal, and the interview-relevant distinction is what each one identifies and returns: remove(value) finds by value and raises ValueError if absent; pop(index) finds by position, defaults to the last element, and is the only one that hands back what it removed; clear() empties the list entirely while keeping the same object identity; and del is a statement, not a method, that can remove a single index or — combined with a slice — an entire range in one step.”

Common Mistakes

  • Calling l.remove(x) expecting it to remove by index — it removes by value; l.remove(2) removes the value 2, not the element at index 2.
  • Not catching ValueError when the value passed to remove() might not be present — check if x in l first, or wrap in try/except.
  • Forgetting pop() with no argument removes the last element, not the first — pop(0) is needed to remove and return the first one, and that call is O(n) (see 10.13 Performance).

10.6 List Operators

+ (Concatenation)

Concatenates two lists into a brand-new list — neither original list is modified.

>>> [1, 2] + [3, 4]
[1, 2, 3, 4]

* (Repetition)

Repeats a list’s elements a given number of times.

>>> [1, 2] * 3
[1, 2, 1, 2, 1, 2]

in / not in

Tests membership — O(n) on a list, since it’s a linear scan (see 10.13 Performance), so convert to a set for repeated checks on large collections.

>>> 2 in [1, 2, 3]
True
>>> 5 not in [1, 2, 3]
True

Comparison

Lists compare element-by-element, lexicographically — the same rule used for tuple and string comparison (see 9.5 String Operators).

>>> [1, 2] == [1, 2]
True
>>> [1, 2] < [1, 3]     # first elements equal, second decides
True

Quick Interview Answer

“+ and * both build a brand-new list rather than mutating either operand — + concatenates, * repeats. in/not in are O(n) linear scans, which matters if membership is being checked repeatedly on a large list; converting to a set first turns that into O(1) average-case lookups. Comparison operators compare element-by-element in order, exactly like lexicographic string comparison — the first differing pair of elements decides the result.”

Common Mistakes

  • Using [[0] * 3] * 3 to build a matrix — * repeats references to the same inner list, so mutating one “row” mutates all of them; a comprehension (see 10.9 Traversing and List Comprehensions) avoids this.
  • Repeatedly checking x in large_list inside a hot loop instead of converting to a set once beforehand, turning an O(n) scan into an O(1) lookup per check.
  • Comparing lists of different lengths and being surprised by the result — [1, 2] < [1, 2, 3] is True because the shorter list is treated as “less than” once it runs out of elements to compare, the same rule as string comparison.

10.7 List Methods

flowchart TD LM["list methods"] LM --> ADD["Add\nappend, extend, insert"] LM --> REM["Remove\nremove, pop, clear"] LM --> QRY["Query\nindex, count"] LM --> RE["Reorder / Copy\nsort, reverse, copy"]

All eleven list methods, grouped by what they do. append/extend/insert and remove/pop/clear were already covered individually in 10.4 Modifying Lists and 10.5 Removing Elements — this is the complete lookup table.

MethodExampleResult
append(x)[1,2].append(3)[1, 2, 3]
extend(it)[1,2].extend([3,4])[1, 2, 3, 4]
insert(i, x)[1,2].insert(0, 0)[0, 1, 2]
remove(x)[1,2,3,2].remove(2)[1, 3, 2]
pop([i])[1,2,3].pop()returns 3
clear()[1,2].clear()[]
index(x)[1,2,3,2].index(2)1
count(x)[1,2,3,2].count(2)2
sort()[3,1,2].sort()in place → [1, 2, 3], returns None
reverse()[1,2,3].reverse()in place → [3, 2, 1]
copy()l.copy()shallow copy, equivalent to l[:]

Query Methods

index() returns the position of the first matching value and raises ValueError if absent — the list equivalent of str.index() (see 9.7 Common String Methods). count() returns how many times a value appears.

>>> [1, 2, 3, 2].index(2)
1
>>> [1, 2, 3, 2].count(2)
2

sort()

Sorts the list in place and returns None — a common gotcha is writing l = l.sort(), which discards the list entirely. See 10.12 Sorting and Searching for sort() vs. sorted().

>>> l = [3, 1, 2]
>>> l.sort()
>>> l
[1, 2, 3]

reverse()

Reverses the list in place — distinct from the l[::-1] slice idiom, which returns a new list instead.

>>> l = [1, 2, 3]
>>> l.reverse()
>>> l
[3, 2, 1]

copy()

Returns a shallow copy — equivalent to l[:] (see 10.3 List Indexing and Slicing and 10.11 Copying and Mutability).

>>> l2 = l.copy()
>>> l2 is l
False

Quick Interview Answer

“The eleven list methods split cleanly into four groups: add (append, extend, insert), remove (remove, pop, clear), query (index, count — read-only, no mutation), and reorder/copy (sort, reverse, copy). The one detail worth memorizing cold: sort() and reverse() both mutate in place and return None — assigning their result (l = l.sort()) silently throws the list away.”

Common Mistakes

  • Writing l = l.sort() expecting the sorted list back — sort() returns None; the list is already sorted in place, so just use l directly afterward.
  • Using index() when the value might not be present — it raises ValueError instead of returning a sentinel like -1; check in first or catch the exception.
  • Confusing l.reverse() (in-place, mutates, returns None) with l[::-1] or reversed(l) (both return a new sequence, original untouched).

10.8 Built-in Functions

These are called as len(l), not l.len() — global functions, not methods, the same distinction covered for strings in 9.8 String Functions.

FunctionExampleResult
len()len([3, 1, 4, 1, 5])5
max()max([3, 1, 4, 1, 5])5
min()min([3, 1, 4, 1, 5])1
sum()sum([3, 1, 4, 1, 5])14
any()any([0, 0, 1])True
all()all([1, 0, 1])False

sorted()

Returns a new sorted list, leaving the original untouched — unlike l.sort(), which sorts in place and returns None (see 10.7 List Methods).

>>> l = [3, 1, 4, 1, 5]
>>> sorted(l)
[1, 1, 3, 4, 5]
>>> sorted(l, reverse=True)
[5, 4, 3, 1, 1]
>>> l          # original is untouched
[3, 1, 4, 1, 5]

any() / all()

any() is True if at least one element is truthy; all() is True only if every element is — both short-circuit, stopping at the first element that decides the answer.

>>> any([0, 0, 1])
True
>>> all([1, 1, 1])
True
>>> all([1, 0, 1])
False

Quick Interview Answer

“These are built-in functions, not list methods — len(l), not l.len(). sum(), max(), min() do the obvious numeric thing. The one worth calling out specifically is sorted() vs. l.sort(): sorted() always returns a brand-new list and leaves the original alone, while sort() mutates in place and returns None — mixing the two up is one of the most common list-related bugs. any()/all() both short-circuit on the first element that decides the result, so they’re O(1) in the best case even on a huge list.”

Common Mistakes

  • Assuming sorted(l) mutates l — it doesn’t; the original list is completely untouched, only the returned value is sorted.
  • Calling sum() on a list containing non-numeric elements (like strings) — it raises TypeError; use "".join(...) for concatenating strings instead.
  • Forgetting any([]) is False and all([]) is True — both edge cases follow directly from their definitions (no element is truthy → any fails; no element violates truthiness → all vacuously holds) but are easy to get backwards from memory.

10.9 Traversing and List Comprehensions

Traversing

for

The standard, most Pythonic way to iterate — no manual index bookkeeping needed.

>>> for x in [1, 2, 3]:
...     print(x, end=" ")
1 2 3

while

Manual index-based iteration — occasionally needed when the index itself must be controlled directly (skipping, jumping), but for is preferred otherwise.

i = 0
lst = [10, 20, 30]
while i < len(lst):
    print(lst[i], end=" ")
    i += 1
# Output: 10 20 30

enumerate()

Yields (index, value) pairs together — the idiomatic way to get both the position and the value without managing a counter manually.

>>> for i, v in enumerate(["a", "b", "c"]):
...     print(i, v)
0 a
1 b
2 c

zip()

Iterates several lists together in parallel, pairing up corresponding elements — stops at the shortest input.

>>> for a, b in zip([1, 2, 3], ["x", "y", "z"]):
...     print(a, b)
1 x
2 y
3 z

List Comprehensions

A compact, single-line syntax for building a new list by transforming and/or filtering an existing iterable: [expression for item in iterable if condition]. More concise, and often faster, than the equivalent explicit for loop with .append() calls.

Basic

>>> squares = [x**2 for x in range(5)]
>>> squares
[0, 1, 4, 9, 16]

Conditional

Adding if filters which items make it into the result.

>>> evens = [x for x in range(10) if x % 2 == 0]
>>> evens
[0, 2, 4, 6, 8]

Nested

A comprehension inside another comprehension — commonly used to build a matrix; see 10.10 Nested Lists and Matrices.

>>> [[x * y for y in range(3)] for x in range(3)]
[[0, 0, 0], [0, 1, 2], [0, 2, 4]]

Multiple Loops

A single comprehension can iterate more than one for clause, producing every combination — equivalent to nested for loops flattened into one expression.

>>> [(x, y) for x in range(2) for y in range(2)]
[(0, 0), (0, 1), (1, 0), (1, 1)]

Quick Interview Answer

“for is the default way to traverse a list; while is reserved for cases needing direct index control. enumerate() gives index-value pairs without a manual counter, and zip() walks multiple lists in parallel, stopping at the shortest. List comprehensions compress the extremely common ’loop + filter + transform + append’ pattern into one expression — [expr for item in iterable if condition] — and are generally both more readable and faster than the equivalent explicit loop, right up until the logic gets complex enough that the comprehension itself becomes hard to read, at which point a full loop is the better choice.”

Common Mistakes

  • Reaching for a while loop with manual indexing when a plain for loop would do the same job more simply and without off-by-one risk.
  • Forgetting zip() silently stops at the shortest input instead of raising an error — mismatched-length inputs can hide a bug rather than surface one.
  • Writing a deeply nested comprehension with multiple conditions that’s harder to read than the equivalent explicit loop — comprehensions are a readability tool first, not a rule to apply unconditionally.

10.10 Nested Lists and Matrices

flowchart TD M["matrix"] --> R0["row 0: [1, 2, 3]"] M --> R1["row 1: [4, 5, 6]"] M --> R2["row 2: [7, 8, 9]"]

matrix[1][2] → 6 — the first index selects the row, the second selects the column within it.

Create

>>> matrix = [[1, 2, 3], [4, 5, 6], [7, 8, 9]]

Access

Chain two indices: the first selects the row (inner list), the second selects the column within it.

>>> matrix[1][2]
6

Update

The same chained-index syntax works as an assignment target too.

>>> matrix[0][0] = 99
>>> matrix
[[99, 2, 3], [4, 5, 6], [7, 8, 9]]

Iterating a Matrix

A plain for loop over the outer list visits each row — the standard pattern for printing or processing a matrix.

>>> for row in matrix:
...     print(row)
[99, 2, 3]
[4, 5, 6]
[7, 8, 9]

Building a Matrix Correctly

A nested list comprehension is the correct way to build a matrix with genuinely independent rows:

>>> [[0 for _ in range(3)] for _ in range(3)]
[[0, 0, 0], [0, 0, 0], [0, 0, 0]]

Quick Interview Answer

“A nested list represents a matrix by making each element of the outer list itself a list — matrix[row][col] chains two indices to reach a cell. The one genuine gotcha is construction: [[0] * 3] * 3 looks right but the outer * repeats references to the exact same inner list three times, so mutating one row mutates all of them. The safe way to build a matrix with independent rows is a nested list comprehension, [[0 for _ in range(3)] for _ in range(3)], which actually creates a fresh inner list on each outer iteration.”

Common Mistakes

  • Building a matrix with [[0] * cols] * rows — all rows are the same shared list object; matrix[0][0] = 1 changes every row, not just the first.
  • Confusing matrix[row][col] order with [col][row] — Python has no built-in convention beyond whatever the code defines; document which axis comes first.
  • Iterating with nested nested-index loops (for i in range(len(matrix)): for j in range(len(matrix[i])):) when a plain for row in matrix: (or enumerate() if the index is genuinely needed) is simpler and less error-prone.

10.11 Copying and Mutability

The general mechanics of assignment, shallow copy, and deep copy are covered in depth in 6.5 Copying Objects — this page focuses on what’s specific to lists: the shortcuts, the function-argument consequence, and the classic mutable-default-argument bug.

Assignment Is Not a Copy

b = a makes both names reference the exact same list object — mutating through either name is visible through both.

>>> a = [1, 2, 3]
>>> b = a
>>> b.append(4)
>>> a, a is b
([1, 2, 3, 4], True)

Shallow Copy Shortcuts

l.copy() and l[:] (see 10.3 List Indexing and Slicing) both create a shallow copy — a new outer list, but nested elements are still shared references.

>>> import copy
>>> original = [1, 2, [3, 4]]
>>> shallow = original.copy()     # equivalent to original[:] or copy.copy(original)
>>> original[2].append(99)
>>> shallow          # sees the change -- the nested list is SHARED
[1, 2, [3, 4, 99]]

Use copy.deepcopy() instead whenever the list contains nested mutable objects and full independence is required — see 6.5 Copying Objects for the complete deep-copy example.

Mutability and Function Arguments

Because a list passed into a function is the same shared object, a function that mutates its parameter mutates the caller’s list too — visible after the call returns.

def modify(lst):
    lst.append("modified")

>>> l = [1, 2]
>>> modify(l)
>>> l          # caller's list WAS changed
[1, 2, 'modified']

The Mutable Default Argument Trap

Using a list as a function’s default argument value is a classic, list-specific bug: the default is created once, at function-definition time, and shared across every call that doesn’t override it.

>>> def add_item(item, items=[]):     # BUG
...     items.append(item)
...     return items
...
>>> add_item("a")
['a']
>>> add_item("b")     # surprise -- 'a' persisted!
['a', 'b']

The fix — the same pattern covered generally in 5.15 Common Mistakes — is to default to None and create a fresh list inside the function body:

def add_item(item, items=None):
    if items is None:
        items = []
    items.append(item)
    return items

Quick Interview Answer

“b = a shares one list — no copy happens at all. l.copy() and l[:] both make a shallow copy: a new outer list, but any nested mutable objects inside are still shared with the original. copy.deepcopy() is the only one that produces a fully independent structure at every level. This same reference-sharing is exactly why a function that mutates a list parameter changes the caller’s list too, and it’s the root cause of the mutable-default-argument bug: a default list is created once at definition time, not fresh per call, so it silently accumulates state across calls unless None is used as the sentinel instead.”

Common Mistakes

  • Assuming l.copy() produces a fully independent list — it’s shallow; nested mutable objects (like inner lists or dicts) are still shared.
  • Writing def f(items=[]) and being surprised the default list accumulates values across unrelated calls — use items=None and create the list inside the function instead.
  • Passing a list into a function and not expecting it to change after the call — if the function shouldn’t mutate the caller’s list, pass a copy explicitly (f(l.copy())) or have the function copy internally.

10.12 Sorting and Searching

sort() vs. sorted()

sort() sorts a list in place, returning None — use when the original order doesn’t need to be preserved. sorted() returns a new sorted list, leaving the original unchanged — use when both orders are needed.

>>> l = [5, 2, 8, 1, 9]
>>> l.sort()
>>> l
[1, 2, 5, 8, 9]

>>> l2 = [5, 2, 8, 1, 9]
>>> sorted(l2)
[1, 2, 5, 8, 9]
>>> l2               # untouched
[5, 2, 8, 1, 9]

Check every element one at a time until a match is found — O(n), but works on any list, sorted or not.

def linear_search(lst, target):
    for i, v in enumerate(lst):
        if v == target:
            return i
    return -1

>>> linear_search([5, 2, 8, 1, 9], 8)
2

Repeatedly halve the search range by comparing against the middle element — O(log n), but requires the list to already be sorted.

def binary_search(sorted_lst, target):
    lo, hi = 0, len(sorted_lst) - 1
    while lo <= hi:
        mid = (lo + hi) // 2
        if sorted_lst[mid] == target:
            return mid
        elif sorted_lst[mid] < target:
            lo = mid + 1
        else:
            hi = mid - 1
    return -1

>>> binary_search([1, 2, 5, 8, 9], 8)
3

Quick Interview Answer

“sort() mutates in place and returns None; sorted() returns a new list and leaves the original alone — mixing these up is a very common bug. For searching: linear search is O(n) but makes no assumptions about order, so it works on any list. Binary search is O(log n) but requires the list to already be sorted — halving the search range on unsorted data gives wrong answers, not just slow ones. The trade-off is real: if a list is searched many times, sorting it once (O(n log n)) to unlock repeated O(log n) binary searches usually pays for itself quickly.”

Common Mistakes

  • Running binary search on an unsorted list — it doesn’t raise an error, it just silently returns wrong results, since the halving logic assumes order that isn’t actually there.
  • Writing l = l.sort() — sort() returns None, so l becomes None afterward; the list was already sorted in place, no reassignment needed.
  • Re-sorting a list before every single search instead of sorting once and reusing it for many binary searches.

10.13 Performance

Time Complexity

OperationComplexityNotes
l[i] (index)O(1)Direct memory offset
l.append(x)O(1) amortizedSpare capacity absorbs most calls (see 10.1 Introduction to Lists)
l.insert(0, x)O(n)Every existing element must shift right
l.pop()O(1)Removing the last element
l.pop(0)O(n)Removing the first element — everything shifts left
x in lO(n)Linear scan
len(l)O(1)Length is cached, not recounted
l.sort()O(n log n)Timsort

Memory

A list’s per-object overhead plus its spare capacity (see 10.1 Introduction to Lists) means it typically uses somewhat more memory than an equivalent tuple — worth considering for very large, unchanging collections (see 5.4 Sequence Data Types).

append() vs. insert()

append() is O(1) amortized because it only ever touches the end. insert() at any position other than the end is O(n), because every following element must shift over. Prefer append() whenever order allows it.

# Fast: O(1) amortized per call
results = []
for item in data:
    results.append(item)

# Slow: O(n) per call -- O(n^2) overall for n items
results = []
for item in data:
    results.insert(0, item)

Quick Interview Answer

“The list operations worth knowing cold: indexing and len() are O(1); append() and pop() (from the end) are O(1) amortized because they only touch one end of the underlying array; insert(0, x) and pop(0) are both O(n) because everything else has to shift; membership testing (in) is an O(n) linear scan; and sort() is O(n log n), using Timsort. The practical consequence that shows up constantly in real code: building a list by repeatedly inserting at the front is O(n²) overall, while appending and reversing (or just appending in the right order to begin with) is O(n).”

Common Mistakes

  • Using l.insert(0, x) in a loop to build a list in reverse order, turning an O(n) job into O(n²) — append in the correct order, or append and reverse once at the end.
  • Checking x in large_list repeatedly in a hot path instead of converting to a set once for O(1) average-case membership tests.
  • Assuming len(l) re-scans the list — it’s O(1), since the length is cached on the object, not recomputed on each call.

10.14 Common Algorithms

Reverse

def reverse_list(l):
    return l[::-1]

>>> reverse_list([1, 2, 3])
[3, 2, 1]

Remove Duplicates (Order-Preserving)

dict.fromkeys() removes duplicates while preserving first-seen order — a plain set() would also dedupe, but loses ordering, the same trade-off covered for strings in 9.9 String Algorithms.

def remove_duplicates(l):
    return list(dict.fromkeys(l))

>>> remove_duplicates([1, 2, 2, 3, 1])
[1, 2, 3]

Find Max / Find Min

>>> max([3, 7, 2]), min([3, 7, 2])
(7, 2)

Merge

def merge_lists(a, b):
    return a + b

>>> merge_lists([1, 2], [3, 4])
[1, 2, 3, 4]

Split into Chunks

Splitting a list into fixed-size chunks — a common batching pattern for processing large collections in manageable pieces.

def split_list(l, n):
    return [l[i:i+n] for i in range(0, len(l), n)]

>>> split_list([1, 2, 3, 4, 5, 6, 7], 3)
[[1, 2, 3], [4, 5, 6], [7]]

Quick Interview Answer

“These algorithms all lean on tools already covered in this chapter rather than hand-rolled loops: l[::-1] for reversal, dict.fromkeys(l) for order-preserving deduplication (a plain set() dedupes but drops order), max()/min() for extremes, + for merging two lists into a new one, and a slicing comprehension — l[i:i+n] stepped by n — for chunking a list into fixed-size batches, which is the standard pattern for processing a huge collection or API payload in manageable pieces.”

Common Mistakes

  • Using a plain set(l) to deduplicate when the original order needs to be preserved — sets don’t preserve insertion order; dict.fromkeys(l) does.
  • Forgetting the last chunk from split_list() may be shorter than n if the list doesn’t divide evenly — code consuming the chunks needs to handle a partial final chunk.
  • Using a + b to merge two very large lists repeatedly in a loop instead of list.extend() or building with a single itertools.chain() pass — each + allocates an entirely new list.

10.15 Lists in DevOps

Lists and File Handling

Lists are the natural in-memory representation of file contents — one element per line or row.

Iterating an open file object yields one line at a time; collecting them into a list, with .strip() to remove trailing newlines, is a standard pattern.

with open("servers.txt") as f:
    lines = [line.strip() for line in f]

>>> lines
['a', 'b', 'c']

Writing is the reverse: looping over a list and writing each item as its own line.

items = ["a", "b", "c"]
with open("output.txt", "w") as f:
    for item in items:
        f.write(item + "\n")

The csv module reads a file into a list of lists (or a list of dicts with DictReader) — one inner list or dict per row.

>>> import csv
>>> with open("data.csv") as f:
...     rows = list(csv.reader(f))
...
>>> rows
[['name', 'age'], ['Alice', '30']]

Server Inventory

The simplest and most common case: a flat list of hostnames to iterate over for health checks, deployments, or config pushes.

>>> servers = ["web01", "web02", "db01"]
>>> for server in servers:
...     print(f"Checking {server}")
Checking web01
Checking web02
Checking db01

Filtering Amazon EC2 Instances

boto3’s describe_instances() returns nested lists of dicts — list comprehensions (see 10.9 Traversing and List Comprehensions) are the idiomatic way to filter them, such as isolating only running instances. Full working scripts for this pattern live in Python Scripting, including the EC2 Untagged Instances Alert.

instances = [
    {"id": "i-001", "state": "running"},
    {"id": "i-002", "state": "stopped"},
]
>>> running = [i["id"] for i in instances if i["state"] == "running"]
>>> running
['i-001']

Docker

A list of running container names, checked with the in operator for quick presence tests.

>>> containers = ["nginx", "redis", "postgres"]
>>> "nginx" in containers
True

Kubernetes

Filtering pod names by prefix (deployment name) using a list comprehension with str.startswith() (see 9.7 Common String Methods).

>>> pods = ["web-abc123", "web-def456", "api-ghi789"]
>>> web_pods = [p for p in pods if p.startswith("web-")]
>>> web_pods
['web-abc123', 'web-def456']

Logs

Filtering a list of log lines down to just the ones matching a severity level — the same comprehension pattern used throughout this section.

>>> log_lines = ["INFO ok", "ERROR fail", "INFO ok"]
>>> errors = [l for l in log_lines if "ERROR" in l]
>>> errors
['ERROR fail']

Packages

Extracting just the package names from a requirements.txt-style list of pinned versions.

>>> packages = ["requests==2.28.0", "flask==2.0.1"]
>>> names = [p.split("==")[0] for p in packages]
>>> names
['requests', 'flask']

Quick Interview Answer

“Lists are the natural shape for anything read line-by-line or row-by-row: log files, CSV rows, server hostnames, boto3 API results. The recurring pattern across almost every DevOps list use case is a filtering comprehension — [x for x in collection if condition] — whether that’s isolating running Amazon EC2 instances by state, matching Kubernetes pod names by deployment prefix, or pulling error lines out of a log. The in operator handles quick presence checks, like confirming an expected Docker container is currently running.”

Common Mistakes

  • Loading an entire multi-gigabyte log file into a list with f.readlines() when a streaming line-by-line for line in f: loop would avoid holding the whole thing in memory at once.
  • Filtering Amazon EC2 or Kubernetes results with a manual loop and .append() instead of the more idiomatic (and often faster) list comprehension.
  • Treating csv.reader()’s output rows as already-typed values — every field comes back as a str, even numeric-looking ones; see 8.1 Introduction to Type Conversion.

10.16 Common Mistakes

Index Errors

Accessing an index that doesn’t exist — especially easy off-by-one mistakes near a list’s boundaries.

>>> l = [1, 2, 3]
>>> l[5]
Traceback (most recent call last):
IndexError: list index out of range

Modifying a List While Iterating Over It

Removing items from a list while iterating over it directly causes elements to be silently skipped, because the indices shift underneath the iterator as items are removed.

# WRONG: mutating the list you're iterating over
l = [1, 2, 3, 4, 5]
for x in l:
    if x % 2 == 0:
        l.remove(x)     # shifts remaining elements -- some get skipped!

# CORRECT: iterate over a COPY, mutate the original
l = [1, 2, 3, 4, 5]
for x in l[:]:
    if x % 2 == 0:
        l.remove(x)

>>> l
[1, 3, 5]

Shared References

Forgetting that b = a shares the same list object rather than copying it — see 10.11 Copying and Mutability for the full explanation and the shallow-vs-deep distinction.

Mutable Default Arguments

Using a list as a function’s default argument value — the default is created once at definition time and shared across every call that doesn’t override it. Fully explained, with the fix, in 10.11 Copying and Mutability and 5.15 Common Mistakes.

Quick Interview Answer

“Four mistakes account for most list bugs: off-by-one IndexErrors near a list’s boundary; mutating a list while iterating over it directly, which silently skips elements because removal shifts every later index backward — the fix is iterating over a copy (for x in l[:]) while mutating the original; forgetting b = a shares one list rather than copying it, so a mutation through b is visible through a too; and the mutable-default-argument trap, where a list default is created once at function-definition time and silently accumulates across calls unless None is used as the sentinel instead.”

Common Mistakes

  • Removing items from a list with l.remove(x) or del l[i] inside a for x in l: loop over that same list — build a new filtered list, use a comprehension, or iterate over l[:] instead.
  • Assuming two separately-created lists that happen to look identical ([1, 2] == [1, 2]) are the same object — == compares content, is compares identity; see 10.6 List Operators.
  • Defining def f(items=[]) and being surprised the default list accumulates values across unrelated calls — default to None and create the list inside the function body.

10.17 Best Practices

Readable Code

Prefer a list comprehension over an equivalent manual for + .append() loop when it stays short and clear — but fall back to a full loop once the logic gets complex enough that the comprehension becomes hard to read. See 10.9 Traversing and List Comprehensions.

Efficient Methods

Use append() over insert(0, ...) when order allows, and convert to a set for repeated membership testing on large collections. See 10.13 Performance.

Choose the Right Structure

Not every ordered collection should be a list — reach for tuple when the contents are fixed, set when uniqueness matters more than order, and dict when items need to be looked up by a key rather than a position. See 5.4 Sequence Data Types and 5.9 Mutable vs Immutable Types.

Quick Interview Answer

“Three habits cover most of what matters: prefer a comprehension over a manual append() loop while it stays readable, but don’t force a genuinely complex transformation into one just to be terse; default to append() over insert(0, ...) and reach for a set when membership is checked repeatedly, since both are about respecting the actual time complexity of list operations; and don’t reflexively reach for a list at all — a tuple for fixed data, a set for uniqueness, or a dict for key-based lookup is often the structurally correct choice.”

Common Mistakes

  • Forcing a multi-condition, multi-step transformation into a single comprehension purely for brevity, producing something harder to read than the equivalent explicit loop.
  • Defaulting to a list for data that’s never going to change, when a tuple documents that intent and is slightly more memory-efficient.
  • Using a list for a large collection that’s checked for membership repeatedly, instead of a set, and paying an unnecessary O(n) cost on every check.

10.18 Interview Questions

Conceptual Questions

  • What’s the difference between a list and a tuple?
  • Why is list.append() O(1) amortized but list.insert(0, x) O(n)?
  • What’s the difference between sort() and sorted()?
  • What’s the difference between a shallow copy and a deep copy of a list?
  • Why shouldn’t a mutable object be used as a function’s default argument value?

Scenario-Based Questions

  • Every even number needs to be removed from a list while iterating — what’s the safe way to do it, and why does the naive approach fail?
  • A function receives a list and appends to it — does the caller see the change? What if it received an int instead?
  • A large list needs fast, repeated membership checks — what would change, and why?

Coding Questions

  • Write a function to remove duplicates from a list while preserving order.
  • Implement binary search on a sorted list without using any built-in search function.
  • Given a list of dicts (like instance records), write a comprehension that extracts IDs matching a condition.
def remove_duplicates(l):
    return list(dict.fromkeys(l))

>>> remove_duplicates([3, 1, 3, 2, 1])
[3, 1, 2]

Quick Interview Answer

“List interview questions cluster around three things: confirming mutability and reference-sharing are understood (list vs. tuple, shallow vs. deep copy, the mutable-default-argument trap), knowing the time complexity trade-offs that drive real design decisions (append vs. insert(0, ...), list membership vs. a set), and classic coding problems — safe removal during iteration, order-preserving deduplication, binary search — that test comfort with slicing, dict.fromkeys(), and index arithmetic rather than memorized syntax.”

Common Mistakes

  • Answering “a list and a tuple are basically the same, just different brackets” without mentioning mutability, hashability, or the performance/memory implications.
  • Solving the “remove while iterating” scenario by suggesting a while loop with manual index tracking instead of the simpler iterate-over-a-copy pattern.
  • Forgetting binary search’s precondition — the list must already be sorted — and describing it as a general-purpose search technique.

10.19 Hands-on Exercises

Inventory Manager

A small class wrapping a list to manage a collection of items with add/remove/list operations.

class InventoryManager:
    def __init__(self):
        self.items = []

    def add(self, item):
        self.items.append(item)

    def remove(self, item):
        if item in self.items:
            self.items.remove(item)

    def list_all(self):
        return self.items

>>> inv = InventoryManager()
>>> inv.add("server01"); inv.add("server02")
>>> inv.remove("server01")
>>> inv.list_all()
['server02']

Todo App

A list of dicts, each representing one task — a common lightweight data model before reaching for a database.

todos = []

def add_todo(task):
    todos.append({"task": task, "done": False})

def complete_todo(index):
    todos[index]["done"] = True

>>> add_todo("Deploy app"); add_todo("Review PR")
>>> complete_todo(0)
>>> todos
[{'task': 'Deploy app', 'done': True}, {'task': 'Review PR', 'done': False}]

Log Analyzer

Counting occurrences of different severity levels across a list of log lines.

def analyze_logs(lines):
    return {
        "errors": sum(1 for l in lines if "ERROR" in l),
        "warnings": sum(1 for l in lines if "WARNING" in l),
    }

>>> analyze_logs(["INFO x", "ERROR y", "WARNING z", "ERROR w"])
{'errors': 2, 'warnings': 1}

CSV Processor

Parsing a CSV blob into a list of dicts, ready for further filtering or aggregation.

import csv, io

def process_csv(text):
    return list(csv.DictReader(io.StringIO(text)))

>>> process_csv("name,age\nAlice,30\nBob,25")
[{'name': 'Alice', 'age': '30'}, {'name': 'Bob', 'age': '25'}]

Mini Projects

Server inventory tool — group a list of Amazon EC2-style instance records by their current state, the core logic behind any fleet-status dashboard:

def group_by_state(instances):
    result = {}
    for inst in instances:
        result.setdefault(inst["state"], []).append(inst["id"])
    return result

>>> group_by_state([
...     {"id": "i-1", "state": "running"},
...     {"id": "i-2", "state": "stopped"},
...     {"id": "i-3", "state": "running"},
... ])
{'running': ['i-1', 'i-3'], 'stopped': ['i-2']}

Docker tracker — combine the container-membership check from 10.15 Lists in DevOps with a list comprehension to report which expected containers are missing from the running list:

expected = ["nginx", "redis", "postgres", "celery"]
running = ["nginx", "redis", "postgres"]
>>> missing = [c for c in expected if c not in running]
>>> missing
['celery']

Backup manager — keep only the N most recent backups, discarding the rest, a common retention-policy script:

def rotate_backups(backups, keep=3):
    return sorted(backups, reverse=True)[:keep]

>>> rotate_backups([
...     "backup_2026-01-01", "backup_2026-01-05",
...     "backup_2026-01-10", "backup_2026-01-15",
... ], keep=2)
['backup_2026-01-15', 'backup_2026-01-10']
  • AWS inventory report — combine the Amazon EC2 filtering pattern from 10.15 Lists in DevOps with the grouping function above to build a full multi-region, multi-state resource report.
  • Kubernetes inventory — combine the pod-filtering pattern from 10.15 Lists in DevOps with count-by-prefix logic to report how many pods each deployment currently has running.

Quick Interview Answer

“These exercises combine the chapter’s tools into small, realistic programs: the inventory manager and todo app wrap a list inside a class or module-level state with add/remove/query operations; the log analyzer and CSV processor lean on comprehensions and the csv module rather than manual parsing; and the mini projects — grouping Amazon EC2 instances by state, diffing expected vs. running Docker containers, rotating backups by keeping the N most recent — are all variations on the same filtering-and-grouping comprehension pattern applied to real infrastructure data.”

Common Mistakes

  • Mutating self.items from outside the InventoryManager class directly instead of going through add()/remove(), bypassing whatever validation those methods might do.
  • Using a list index directly as a todo item’s permanent identifier (complete_todo(index)) — removing an earlier item shifts every later index, silently completing the wrong task.
  • Sorting backup filenames as plain strings and assuming that sorts chronologically — it only works if the naming format is zero-padded and lexicographically ordered the same as chronological order, as in the ISO-style dates used above.

Lists: Chapter Practice

Mini lab

Filter unhealthy server names into a new list.

Expected output: [‘web-01’, ‘web-02’]

HintUse the chapter examples to choose the operation, then run your version before opening the solution.
Show solution
names = ["web-01", "db-01", "web-02"]
print([name for name in names if name.startswith("web-")])

Knowledge check

What does list.append() return?

Check your answerNone; it mutates the existing list.

Finish the chapter

  • Explain each operation without reading the lesson.
  • Change the sample input and predict the output before running it.
  • Revisit the chapter if the result differs from your prediction.

Back to this chapter · Choose the next chapter

11.1 Introduction to Tuples

Python Tuples

What Is a Tuple?

What Is It?

An ordered, immutable collection — like a list, but once created it can never be changed. Written with parentheses (optional, but conventional), comma-separated.

>>> point = (3, 4)
>>> point
(3, 4)

Why Does It Matter?

A tuple represents data that’s a fixed record or grouping that should never be accidentally modified — the opposite design goal from Chapter 10: Lists.

Characteristics

Why Use Tuples?

Immutability is a feature, not a limitation: it guarantees a value can’t be accidentally changed elsewhere in a program, makes it safely shareable, and is exactly what’s required to use a compound value as a dict key.

Real-World Applications

  • A coordinate pair (x, y)
  • A database row/record returned from a query
  • RGB color values (255, 0, 0)
  • Function return values that bundle multiple results together

Tuples in DevOps

Fixed configuration values that should never change at runtime — AWS regions, allowed ports, server details — are naturally represented as tuples. 11.12 Tuples in DevOps is dedicated entirely to these patterns.

>>> DB_CONFIG = ("localhost", 5432, "mydb")

Internal Representation

flowchart LR T["tuple (1, 2, 3)"] --> M1["exact-size array\nno spare capacity"] L["list [1, 2, 3]"] --> M2["over-allocated array\nspare capacity reserved"]

A tuple’s size is known and fixed at creation, so CPython allocates exactly enough space — nothing extra.

Unlike a list (see 10.1 Introduction to Lists), a tuple’s exact size is known at creation time, so CPython allocates exactly enough space for its elements — no spare capacity reserved, because it can never grow.

>>> import sys
>>> sys.getsizeof((1, 2, 3))
64
>>> sys.getsizeof([1, 2, 3])     # same data, but list reserves extra room
88

Once built, a tuple’s contents — which object each slot references — can never be reassigned; attempting to do so raises TypeError.

>>> t = (1, 2, 3)
>>> t[0] = 99
Traceback (most recent call last):
TypeError: 'tuple' object does not support item assignment

Like a list, each tuple slot holds a reference to an object, not the object’s data directly — this is exactly what allows a tuple to contain a mutable object even though the tuple itself is immutable (see 11.8 Nested Tuples).

No spare capacity to manage, no possibility of resizing, and (for tuples of hashable items) a cached hash — all of which make tuples marginally faster to create and smaller in memory than an equivalent list; full comparison in 11.10 Performance.

Quick Interview Answer

“A tuple is an ordered, immutable collection — everything a list is, minus the ability to change after creation. That immutability isn’t a limitation, it’s the entire point: it makes a tuple safe to share without defensive copying, and it’s exactly what makes a tuple hashable and therefore usable as a dict key or set member, which a list can never be. Internally, because a tuple’s size is fixed and known at creation, CPython allocates exactly enough memory for it — no spare capacity the way a list reserves for future append() calls — which is why a tuple is both smaller and marginally faster to create than an equivalent list.”

Common Mistakes

  • Treating a tuple as “just an immutable list” without connecting why that matters — the practical payoff is hashability (dict keys, set members) and safety from accidental mutation, not just a syntax restriction.
  • Assuming a tuple reserves spare capacity the way a list does — it doesn’t, and can’t, since it’s never going to grow.
  • Reaching for a list by default even when the data is a fixed, unchanging record — a tuple documents that intent and is slightly cheaper.

11.2 Creating Tuples

Empty Tuple

>>> t = ()
>>> t
()

Single-Element Tuple

A tuple with exactly one element requires a trailing comma — (1) is not a tuple, it’s just the int 1 wrapped in redundant parentheses. This is one of the most common tuple mistakes; see 11.13 Common Mistakes.

>>> single = (1,)          # the comma is what makes it a tuple
>>> type(single)
<class 'tuple'>
>>> not_a_tuple = (1)      # no comma -- this is just an int
>>> type(not_a_tuple)
<class 'int'>

Tuple Literals

>>> coordinates = (3, 4, 5)

tuple() Constructor

Converts any iterable into a tuple — covered in depth in 8.6 Collection Type Conversion.

>>> tuple([1, 2, 3])
(1, 2, 3)
>>> tuple("abc")
('a', 'b', 'c')

Nested Tuples

>>> nested = ((1, 2), (3, 4))

Mixed Data Types

>>> record = (1, "two", 3.0, True)

Quick Interview Answer

“A tuple literal is just comma-separated values, parentheses optional except for one critical case: a single-element tuple requires a trailing comma — (1,) is a tuple, (1) is just the int 1, since parentheses alone are only grouping, not tuple syntax. tuple() converts any iterable, the same as list(), including splitting a string into its individual characters. Tuples nest and mix types exactly like lists do, since every slot is just a reference regardless of what it points to.”

Common Mistakes

  • Writing (1) when a one-element tuple was intended — it silently produces an int, not a tuple, with no error to flag the mistake.
  • Forgetting that the comma, not the parentheses, is what actually creates a tuple — 1, 2, 3 (no parentheses at all) is still a valid tuple literal.
  • Calling tuple("hello") expecting a single-element tuple containing the whole string — it produces one element per character, the same behavior as list() on a string.

11.3 Tuple Indexing and Slicing

Identical indexing and slicing rules to lists (see 10.3 List Indexing and Slicing) and strings — only the container type, and one copying detail at the end, differs.

Indexing

>>> t = (10, 20, 30, 40, 50)
>>> t[0], t[2]
(10, 30)
>>> t[-1]
50

Nested Tuple Indexing

>>> nested = ((1, 2), (3, 4))
>>> nested[0][1]
2

IndexError

>>> t[10]
Traceback (most recent call last):
IndexError: tuple index out of range

Slicing

>>> t = (10, 20, 30, 40, 50)
>>> t[1:3]
(20, 30)
>>> t[:2], t[2:], t[::2]
((10, 20), (30, 40, 50), (10, 30, 50))
>>> t[::-1]      # reverse
(50, 40, 30, 20, 10)

Copying Tuples

What’s different from lists: because tuples are immutable, CPython optimizes t[:] to return the same object rather than building a new one — there’s no risk in sharing it, since neither copy can ever be mutated. Contrast with 10.3 List Indexing and Slicing, where l[:] on a list always creates a new object.

>>> copy_t = t[:]
>>> copy_t is t     # SAME object -- safe because tuples can't be mutated
True

Quick Interview Answer

“Indexing and slicing on a tuple work exactly like lists and strings — same t[start:stop:step] syntax, direct indexing raises IndexError out of range, slicing never does. The one detail that’s genuinely different from lists: t[:] doesn’t copy anything at all — CPython just hands back the same object, because there’s zero risk in two names sharing an object that can never be mutated. That’s the same optimization strings get for the same reason, and the exact opposite of what happens with a list slice, which always allocates a new list.”

Common Mistakes

  • Expecting t[:] on a tuple to behave like l[:] on a list and produce a distinct copy — is returns True for the tuple case, False for the list case, and that difference is a direct, testable consequence of immutability.
  • Forgetting slicing a tuple returns another tuple, not a list — t[1:3] on (10, 20, 30) gives (20, 30), parentheses and all.
  • Assuming a negative-step slice mutates or reorders the original tuple — t[::-1] returns a brand-new tuple; t itself is untouched, since it couldn’t be touched even if that were the intent.

11.4 Tuple Packing and Unpacking

flowchart LR P1["1"] --> PK["packing"] P2["2"] --> PK P3["3"] --> PK PK --> PT["point = (1, 2, 3)"] PT --> UP["unpacking"] UP --> U1["x = 1"] UP --> U2["y = 2"] UP --> U3["z = 3"]

Packing bundles values into a tuple; unpacking spreads them back out.

Packing

Writing several comma-separated values (parentheses optional) automatically bundles them into a single tuple.

>>> point = 1, 2, 3     # parentheses are optional -- this IS a tuple
>>> point
(1, 2, 3)

Unpacking

The reverse: assigning a tuple to several comma-separated names distributes its values into them, one per name.

>>> x, y, z = point
>>> x, y, z
(1, 2, 3)

Multiple Assignment and the Swap Idiom

Multiple assignment (a, b = 1, 2) is really packing and unpacking happening together in a single statement — this is also what makes the classic swap idiom work without a temporary variable.

>>> a, b = 10, 20
>>> a, b = b, a     # swap -- no temp variable needed
>>> a, b
(20, 10)

Extended Unpacking

The * operator captures “everything else” into a list during unpacking — it can be placed first, last, or in the middle.

>>> first, *rest = (1, 2, 3, 4, 5)
>>> first, rest
(1, [2, 3, 4, 5])

>>> *init, last = (1, 2, 3, 4, 5)
>>> init, last
([1, 2, 3, 4], 5)

>>> a, *mid, z = (1, 2, 3, 4, 5)
>>> a, mid, z
(1, [2, 3, 4], 5)

Quick Interview Answer

“Packing is comma-separated values auto-bundling into a tuple — parentheses are cosmetic, the comma is what matters. Unpacking is the reverse: assigning a tuple to several comma-separated names. Multiple assignment is just packing the right side and unpacking it into the left side in one statement, which is exactly what makes a, b = b, a work as a swap with no temporary variable — the right side is fully packed into a temporary tuple before any name on the left gets rebound. Extended unpacking with *name captures ’everything else’ into a list, and can appear anywhere in the unpacking target — first, last, or in the middle.”

Common Mistakes

  • Assuming a, b = b, a evaluates left-to-right like separate statements — the entire right-hand side is packed into a tuple first, before any assignment happens, which is exactly why it swaps correctly instead of overwriting a before b reads it.
  • Unpacking a tuple into the wrong number of names — x, y = (1, 2, 3) raises ValueError: too many values to unpack, unless a * catch-all is used to absorb the extra values.
  • Forgetting extended unpacking’s *rest always produces a list, not a tuple, even though the source was a tuple.

11.5 Tuple Operators

+ (Concatenation)

Combines two tuples into a new tuple — since tuples are immutable, this can never modify either original.

>>> (1, 2) + (3, 4)
(1, 2, 3, 4)

* (Repetition)

>>> (1, 2) * 3
(1, 2, 1, 2, 1, 2)

Membership (in, not in)

>>> 2 in (1, 2, 3), 5 not in (1, 2, 3)
(True, True)

Comparison Operators

Tuples compare element-by-element, lexicographically — the same rule that applies to lists and strings (see 10.6 List Operators and 9.5 String Operators), and exactly the mechanism behind version-tuple comparison like (1, 2, 0) < (1, 3, 0).

>>> (1, 2) == (1, 2)
True
>>> (1, 2) < (1, 3)
True

Quick Interview Answer

“Tuple operators mirror list operators exactly, with the one obvious difference: + and * are the only ways to build a bigger tuple from a smaller one, since there’s no in-place append() or extend() to reach for instead — both always allocate a brand-new tuple. Comparison is lexicographic, element-by-element, the same rule used for strings and lists — which is also exactly what makes comparing version tuples like (1, 2, 0) < (1, 10, 0) work correctly, unlike comparing the equivalent strings '1.2.0' < '1.10.0'.”

Common Mistakes

  • Reaching for t.append(x) on a tuple out of habit — tuples have no such method; t = t + (x,) (or t = (*t, x)) is the way to “add” an element, producing a new tuple each time.
  • Building up a large tuple incrementally with repeated += in a loop — just like the equivalent list mistake, each concatenation allocates a new tuple and copies everything so far, an O(n²) pattern for n pieces.
  • Comparing a tuple of strings and a tuple of numbers expecting a meaningful result — mixed-type element comparison raises TypeError in Python 3, the same as comparing a bare str to an int directly.

11.6 Tuple Methods

Tuples have only two methods — a direct consequence of immutability. Every method that would modify a list (append, remove, sort, …; see 10.7 List Methods) simply doesn’t exist for tuples, since there’s nothing for it to do.

count()

Counts how many times a value appears.

>>> t = (1, 2, 3, 2, 1)
>>> t.count(2)
2

index()

Returns the position of the first matching value; raises ValueError if absent — the same contract as list.index().

>>> t.index(2)
1

Quick Interview Answer

“A tuple has exactly two methods, count() and index() — both read-only. Every list method that mutates (append, insert, remove, sort, reverse, …) has no tuple equivalent at all, not because it was left out, but because immutability makes it structurally meaningless — there’s no operation for a method called append to perform on an object that can never grow. This short method list is itself one of the clearest signals of what immutability actually costs and buys: fewer things you can do, in exchange for the guarantees covered in 11.9 Immutability and Copying.”

Common Mistakes

  • Calling t.append(x), t.remove(x), or t.sort() on a tuple — none exist; Python raises AttributeError: 'tuple' object has no attribute 'append' (and so on) rather than silently doing nothing.
  • Using index() when the value might not be present — it raises ValueError, same as list.index(); check in first or catch the exception.
  • Expecting count() or index() to accept a start/stop range the way str.find() does — tuple.index() does actually accept optional start/stop arguments, but it’s easy to forget since count() does not.

11.7 Built-in Functions and Traversal

Built-in Functions

The same general-purpose sequence functions that work on lists (see 10.8 Built-in Functions) work identically on tuples.

>>> len((3, 1, 4, 1, 5))
5
>>> max((3, 1, 4, 1, 5)), min((3, 1, 4, 1, 5))
(5, 1)
>>> sum((3, 1, 4, 1, 5))
14
>>> any((0, 0, 1)), all((1, 1, 1))
(True, True)

sorted()

Always returns a list, even when given a tuple as input — there’s no in-place tuple.sort(), since tuples can’t be modified.

>>> sorted((3, 1, 4, 1, 5))
[1, 1, 3, 4, 5]

Traversing Tuples

for Loop

>>> for x in (1, 2, 3):
...     print(x, end=" ")
1 2 3

while Loop

i = 0
t = (10, 20, 30)
while i < len(t):
    print(t[i], end=" ")
    i += 1
# Output: 10 20 30

enumerate()

>>> for i, v in enumerate(("a", "b", "c")):
...     print(i, v)
0 a
1 b
2 c

Quick Interview Answer

“Every general-purpose sequence function — len, max, min, sum, any, all — works identically on a tuple as it does on a list, since they only rely on the iterable and comparison protocols, not mutability. The one worth calling out specifically is sorted(): it always returns a list, regardless of what type went in, because there’s no such thing as an in-place tuple.sort() — sorting a tuple’s elements necessarily means producing a different object. Traversal is identical to lists too: for is the default, while for manual index control, enumerate() for index-value pairs without a counter.”

Common Mistakes

  • Expecting sorted(some_tuple) to return a tuple back — it always returns a list; wrap in tuple(sorted(...)) if a tuple result is actually needed.
  • Calling some_tuple.sort() expecting in-place sorting like a list — tuples have no sort() method at all (see 11.6 Tuple Methods).
  • Using a while loop with manual indexing when a plain for loop accomplishes the same traversal more simply.

11.8 Nested Tuples

Creating and Accessing

>>> nested = ((1, 2, 3), (4, 5, 6))
>>> nested[1][2]
6

Updating Nested Mutable Objects

A tuple’s immutability only applies to its own slots — if one of those slots holds a mutable object (like a list), that inner object can still be freely modified. The tuple itself never changes (it still references the same list object); only the list’s own contents change.

flowchart LR T["tuple (immutable container)"] --> S0["t[0] = 1"] T --> S1["t[1] = 2"] T --> S2["t[2] -> list"] S2 --> L["[3, 4] (mutable)"]

t[2].append(99) is allowed — it mutates the inner list. t[2] = [9, 9] is not — it would rebind a tuple slot.

>>> t = (1, 2, [3, 4])
>>> t[2].append(99)      # ALLOWED -- mutates the inner list, not the tuple
>>> t
(1, 2, [3, 4, 99])
>>> t[2] = [1]            # NOT allowed -- would rebind a tuple slot
Traceback (most recent call last):
TypeError: 'tuple' object does not support item assignment

This same shallow-immutability principle is exactly why a tuple containing a list is not hashable — see 11.9 Immutability and Copying.

Quick Interview Answer

“A tuple’s immutability protects its own slots — which object each position refers to — but says nothing about what that object does internally. t = (1, 2, [3, 4]) can never have t[2] reassigned to a different object, but the list at t[2] is still just an ordinary mutable list, and t[2].append(99) works fine. This is the single most misunderstood thing about tuples: ‘immutable’ means the container’s own bindings are frozen, not that everything reachable through it is frozen too — and it’s exactly why a tuple holding a list can’t be hashed, since its effective contents could still silently change out from under a dict or set that was keying on it.”

Common Mistakes

  • Describing a tuple as “fully immutable” without the shallow-immutability caveat — it’s a very common wrong answer to “can a tuple ever change?”
  • Assuming t[2].append(99) is disallowed because it “modifies the tuple” — it doesn’t modify the tuple at all; the tuple still points at the exact same list object before and after.
  • Trying to hash a tuple that contains a list and being surprised by TypeError: unhashable type: 'list' — the tuple’s own hashability depends entirely on whether every element it holds is itself hashable.

11.9 Immutability and Copying

The general mutable-vs-immutable distinction is covered in 5.9 Mutable vs Immutable Types, and general copying mechanics in 6.5 Copying Objects. This page focuses on what immutability specifically buys tuples, and the one copying behavior that’s genuinely different from every mutable type.

Why Tuples Are Immutable

By design: a tuple represents a fixed, complete record — immutability is what guarantees that record can never be silently altered by code elsewhere that happens to hold a reference to it.

Benefits

  • Safe to share across functions/threads without defensive copying.
  • Hashable (if contents are hashable) — usable as a dict key or set member.
  • Communicates intent — signals to readers that this data is a fixed record, not a growing collection.
>>> t1 = (1, 2, 3)
>>> hash(t1)                # tuples of hashable items ARE hashable
529344067295497451
>>> d = {t1: "value"}       # usable as a dict key
>>> d
{(1, 2, 3): 'value'}

Limitations

A tuple can’t grow, shrink, or reorder in place — any “modification” requires building an entirely new tuple. A tuple containing a mutable element (like a list) is also not hashable, since its effective contents could still change (see 11.8 Nested Tuples).

>>> hash((1, [2, 3]))       # contains a list -- unhashable
Traceback (most recent call last):
TypeError: unhashable type: 'list'

Copying Tuples

b = a makes b reference the exact same tuple object as a — ordinary reference assignment, same as any other type.

>>> a = (1, 2, 3)
>>> b = a
>>> a is b
True

What’s different from lists: because a tuple can never be mutated, sharing a reference to it — via assignment or copy.copy() — is always completely safe. There’s no risk of one name’s changes surprising another, since neither can change anything.

>>> import copy
>>> c = copy.copy(a)
>>> c is a     # even an explicit copy returns the SAME object -- safe due to immutability
True

copy.deepcopy() still recursively copies a tuple that contains nested mutable objects, exactly as described in 6.5 Copying Objects — the shallow-copy optimization above only kicks in because a plain shallow copy of an immutable container has nothing to protect against.

Quick Interview Answer

“Immutability buys a tuple three concrete things: it’s safe to share across functions and threads with zero defensive copying, it’s hashable — as long as every element inside it is also hashable — which makes it usable as a dict key or set member, and it signals to a reader that this data is a fixed record. The copying behavior follows directly: since a tuple can never be mutated, even copy.copy() just hands back the identical object instead of duplicating anything, because there’s nothing to protect against by copying. copy.deepcopy() still does real work if the tuple holds nested mutable objects, since those inner objects genuinely could be mutated.”

Common Mistakes

  • Assuming any tuple is automatically hashable — it’s only hashable if every element inside it is hashable too; a tuple containing a list raises TypeError: unhashable type.
  • Expecting copy.copy() on a tuple to produce a genuinely separate object the way it does for a list — for an immutable object, CPython optimizes this to return the same object, which is observably correct behavior, not a bug.
  • Forgetting copy.deepcopy() still matters for a tuple that holds mutable elements — immutability at the tuple’s own level doesn’t make everything inside it independent.

11.10 Performance

flowchart TD A["Memory (3 items)"] --> A1["tuple: 64 bytes"] A --> A2["list: 88 bytes"] B["Literal creation speed"] --> B1["tuple: faster"] B --> B2["list: slower"] C["Why"] --> C1["tuple: no over-allocation needed"] C --> C2["list: reserves extra capacity for append()"]

Tuples skip the extra bookkeeping lists need to support mutation — the right choice for fixed, read-only data.

Memory Usage

>>> import sys
>>> sys.getsizeof((1, 2, 3))
64
>>> sys.getsizeof([1, 2, 3])
88

Tuple vs List

Propertytuplelist
Mutable?NoYes
Memory (3 items)64 bytes88 bytes
Hashable?Yes (if contents are)Never
Has spare capacity?NoYes
Methods available2 (count, index)11 (see 10.7 List Methods)

Speed Comparison

Creating a tuple literal is measurably faster than creating an equivalent list literal, since there’s no spare-capacity bookkeeping to set up:

>>> import timeit
>>> timeit.timeit(lambda: (1, 2, 3, 4, 5), number=1000000)
0.0246     # seconds -- illustrative, will vary by machine
>>> timeit.timeit(lambda: [1, 2, 3, 4, 5], number=1000000)
0.0623

Quick Interview Answer

“Tuples are both smaller and faster than an equivalent list, and both differences trace back to the same root cause: a list has to reserve spare capacity and manage the possibility of growing, while a tuple’s size is fixed forever at creation, so CPython allocates exactly what’s needed and skips all of that bookkeeping. The trade-off is real, though — a tuple gives up 9 of the 11 list methods (only count() and index() remain) in exchange for that speed and the hashability that comes with immutability. For genuinely fixed data, that trade is a clear win; for anything that grows or shrinks, it isn’t a choice at all — a list is required.”

Common Mistakes

  • Treating the tuple/list performance gap as large enough to matter for ordinary application code — it’s real but small; the decision between them should be driven by mutability and hashability needs, not micro-benchmarks.
  • Converting a list to a tuple purely for a speed boost in a hot loop where the collection still needs to grow or shrink — that requirement rules out a tuple regardless of performance.
  • Forgetting the memory difference scales with collection size — the gap in a small tuple/list pair is a few dozen bytes, but the same relative saving matters more when creating millions of small fixed records.

11.11 Common Algorithms

Search and aggregate operations on a tuple use the exact same techniques as on a list (see 10.12 Sorting and Searching and 10.14 Common Algorithms) — the only difference is a tuple can’t be sorted in place.

Searching

def search(tup, target):
    for i, v in enumerate(tup):
        if v == target:
            return i
    return -1

>>> search((5, 2, 8, 1, 9), 8)
2

Counting

>>> (5, 2, 8, 1, 9).count(2)
1

Finding Maximum / Minimum

>>> max((5, 2, 8, 1, 9)), min((5, 2, 8, 1, 9))
(9, 1)

Quick Interview Answer

“Every list algorithm — linear search, counting, finding extremes — applies to a tuple unchanged, since none of them require mutation, only iteration and comparison. The one thing that genuinely doesn’t carry over is in-place sorting: sorted(some_tuple) works fine and returns a list, but there’s no tuple.sort(), because sorting in place is a contradiction for something that can’t be modified. If a sorted tuple is actually needed, wrap the result: tuple(sorted(t)).”

Common Mistakes

  • Reaching for t.sort() on a tuple, forgetting it doesn’t exist — use sorted(t) (returns a list) or tuple(sorted(t)) if a tuple result is needed.
  • Writing a custom linear search for a tuple instead of just using in or .index() when only presence or a single position is needed, not every algorithmic detail of the search.
  • Assuming binary search needs adapting for tuples — it works identically, since it only relies on indexing and comparison, neither of which differs from a list.

11.12 Tuples in DevOps

Tuples and File Handling

Files store plain text, so “reading a tuple” means parsing each line and explicitly reconstructing the tuple — there’s no native tuple format on disk.

read_back = []
with open("records.txt") as f:
    for line in f:
        name, age = line.strip().split(",")
        read_back.append((name, int(age)))

>>> read_back
[('Alice', 30), ('Bob', 25)]

Writing is the reverse: unpack each tuple’s fields directly in the write loop.

records = [("Alice", 30), ("Bob", 25)]
with open("records.txt", "w") as f:
    for name, age in records:
        f.write(f"{name},{age}\n")

csv.writer accepts any iterable per row, including tuples directly — and csv.reader’s rows can be converted to tuples with tuple(row) if immutability is wanted downstream.

>>> import csv
>>> with open("data.csv") as f:
...     rows = [tuple(row) for row in csv.reader(f)]
...
>>> rows
[('Alice', '30'), ('Bob', '25')]

Five Places Tuples Are the Natural Fit

Specifically because the data they hold should never accidentally change.

Configuration Data

Fixed connection settings, unpacked directly into named variables.

>>> DB_CONFIG = ("localhost", 5432, "mydb")
>>> host, port, dbname = DB_CONFIG
>>> host, port, dbname
('localhost', 5432, 'mydb')

AWS Regions

A fixed allowlist of valid regions — a tuple communicates “this list should never change at runtime” more clearly than a list would.

>>> AWS_REGIONS = ("us-east-1", "us-west-2", "eu-west-1")
>>> "us-east-1" in AWS_REGIONS
True

Port Numbers

>>> HTTP_PORTS = (80, 443, 8080)

Server Details

A single server’s fixed attributes, bundled and unpacked as one record.

>>> server = ("web01", "10.0.1.5", "running")
>>> name, ip, status = server
>>> name, ip, status
('web01', '10.0.1.5', 'running')

Log Records

A single parsed log entry as an immutable record, safe to pass around without worrying about accidental modification — the same shape used for 9.12 Strings in DevOps and AWS’s regex capture groups.

>>> log_record = ("2026-07-13", "10:22:05", "ERROR", "Connection refused")
>>> date, time, level, msg = log_record
>>> level, msg
('ERROR', 'Connection refused')

Quick Interview Answer

“Tuples show up in DevOps code anywhere a value is fixed by design and shouldn’t be accidentally mutated later in the script — configuration bundles unpacked into named variables, an allowlist of valid AWS regions or ports, a single server’s fixed attributes, or a parsed log line’s fields. The common thread is unpacking: host, port, dbname = DB_CONFIG is both more readable and safer than indexing into an unnamed sequence with config[0], config[1], since it documents what each field means at the point of use.”

Common Mistakes

  • Using a list for a fixed allowlist (valid regions, valid ports) when a tuple would both perform slightly better and communicate that the collection is not meant to change.
  • Indexing into a config or record tuple positionally (server[0], server[2]) throughout a codebase instead of unpacking once into named variables — harder to read and fragile if the tuple’s shape ever changes.
  • Forgetting every field parsed from a file or CSV row starts as str, even inside a tuple — the age in ("Alice", "30") needs explicit conversion (see 8.1 Introduction to Type Conversion) before arithmetic.

11.13 Common Mistakes

Single-Element Tuple

Forgetting the trailing comma — (1) is an int, not a tuple (see 11.2 Creating Tuples). This silently produces the wrong type instead of raising an error, making it an easy bug to miss.

>>> not_a_tuple = (1)
>>> type(not_a_tuple)
<class 'int'>
>>> actual_tuple = (1,)     # comma required
>>> type(actual_tuple)
<class 'tuple'>

Modifying Tuples

Attempting to assign to an index, as if it were a list — always raises TypeError.

>>> t = (1, 2, 3)
>>> t[0] = 99
Traceback (most recent call last):
TypeError: 'tuple' object does not support item assignment

Nested Mutable Objects

Assuming a tuple is fully immutable, then being surprised that a nested list inside it can still be modified (see 11.8 Nested Tuples) — the tuple’s immutability is shallow, not deep.

>>> t = (1, [2, 3])
>>> t[1].append(4)     # allowed -- the inner list is still mutable
>>> t
(1, [2, 3, 4])

Quick Interview Answer

“Three mistakes account for most tuple bugs: forgetting the trailing comma on a single-element tuple, which silently produces the wrong type instead of erroring; trying to assign to an index like it’s a list, which correctly raises TypeError — that one’s rarely a silent bug, just a stumbling block for people newer to the type; and assuming immutability is deep, then being confused when a nested list inside a tuple changes anyway. That last one is the most conceptually important, since it reveals what ‘immutable’ actually means for a container type — the slots are frozen, not everything reachable through them.”

Common Mistakes

  • Writing config = (value) intending a one-element tuple and later iterating over it, hitting a TypeError because config turned out to be whatever type value already was.
  • Trying t.remove(x) or t[i] = x on a tuple used for config data, discovering only at runtime that the “list-like” object doesn’t support the mutation.
  • Passing a tuple containing a list to something expecting a hashable key (a dict key, a set element) and hitting TypeError: unhashable type: 'list' instead of catching it during design.

11.14 Best Practices

When to Use Tuples

  • The data represents a fixed record (coordinates, RGB values, a database row).
  • The value needs to be a dict key or set member.
  • The intent is to signal to readers that this collection should never change.

Efficient Coding

Use tuple unpacking to make multi-value returns and record access self-documenting, instead of indexing into an unnamed sequence.

# Less readable
server = ("web01", "10.0.1.5", "running")
print(server[0], server[2])

# More readable -- unpack into named variables
name, ip, status = server
print(name, status)

Choosing Tuple vs. List

SituationBest choice
Fixed record that shouldn’t changetuple
Collection that will grow/shrink/reorderlist
Needs to be a dict key or set membertuple
Returning multiple values from a functiontuple

Quick Interview Answer

“The decision between tuple and list almost always comes down to one question: does this collection’s size or contents ever need to change after creation? If not — a coordinate, an RGB value, a config record, a function’s multi-value return — a tuple is the right default, both for the small performance win and, more importantly, because it documents that intent to every future reader. The other habit worth building is unpacking over indexing: name, ip, status = server reads at a glance; server[0], server[2] scattered through a codebase doesn’t, and breaks silently if the tuple’s shape ever changes.”

Common Mistakes

  • Defaulting to a list everywhere out of habit, even for data that’s structurally a fixed record and never mutates.
  • Indexing into a tuple positionally throughout a codebase instead of unpacking once into named variables at the point of receipt.
  • Choosing a tuple for a collection that actually does need to grow later, then working around the immutability with awkward t = t + (x,) patterns instead of just using a list from the start.

11.15 Interview Questions

Conceptual Questions

  • What’s the difference between a tuple and a list?
  • Why does (1) not create a tuple, but (1,) does?
  • Why are tuples hashable but lists are not?
  • Can a tuple contain mutable objects? What are the implications?
  • Why is tuple creation generally faster than list creation?

Scenario-Based Questions

  • A fixed set of valid AWS regions should never be modified at runtime — would you use a tuple or a list, and why?
  • A tuple contains a list as one of its elements — is the tuple hashable? Why or why not?
  • Three values need to be returned from a function — what’s the idiomatic way to do this in Python, and how would the caller consume it?

Coding Questions

  • Write a function that swaps two variables’ values using tuple packing/unpacking, without a temporary variable.
  • Given a tuple of mixed types, write code that safely determines whether it’s hashable.
  • Write a function returning both the min and max of a tuple in a single call, using multiple return values.
def minmax(t):
    return min(t), max(t)

>>> minmax((5, 2, 8, 1))
(1, 8)

Quick Interview Answer

“Tuple questions cluster around three things: immutability and its consequences (hashability, safety from accidental change, why (1) isn’t a tuple but (1,) is), the tuple-vs-list decision (fixed record vs. growable collection, dict-key eligibility), and the packing/unpacking idioms that show up constantly in real code — the temp-variable-free swap, and returning multiple values from a function as an implicit tuple. The hashability scenario question is a good one to prepare precisely: a tuple is hashable only if every element inside it is also hashable, so a tuple containing a list is not.”

Common Mistakes

  • Answering “tuples and lists are basically interchangeable, just pick either” without mentioning mutability, hashability, or when each is structurally the correct choice.
  • Solving the swap question by declaring a temporary variable — missing the entire point that a, b = b, a packs the right side into a tuple before any assignment happens.
  • Forgetting hashability is contents-dependent, not automatic for every tuple — answering “yes, all tuples are hashable” without the caveat about nested mutable elements.

11.16 Hands-on Exercises

Tuple Operations

Practice combining concatenation, membership, and slicing on a single tuple.

>>> t = (10, 20, 30)
>>> t = t + (40,)     # "append" by concatenating a new tuple
>>> t
(10, 20, 30, 40)
>>> 20 in t
True

Packing & Unpacking

A swap function is the classic demonstration of packing and unpacking working together.

def swap(a, b):
    return b, a

>>> swap(1, 2)
(2, 1)

Data Processing

Returning multiple related results from one function call, using a tuple as the return type.

def minmax(t):
    return min(t), max(t)

>>> minmax((5, 2, 8, 1))
(1, 8)

Mini Projects

Server configuration — a fixed server record, unpacked wherever it’s used; the tuple guarantees the config can’t be accidentally mutated mid-script:

SERVER = ("web01", "10.0.1.5", 8080, "running")

def describe(server):
    name, ip, port, status = server
    return f"{name} ({ip}:{port}) is {status}"

>>> describe(SERVER)
'web01 (10.0.1.5:8080) is running'

AWS region mapper — mapping a fixed tuple of regions to their index, useful for consistent ordering or round-robin selection logic:

def region_mapper(regions):
    return {r: i for i, r in enumerate(regions)}

>>> region_mapper(("us-east-1", "us-west-2", "eu-west-1"))
{'us-east-1': 0, 'us-west-2': 1, 'eu-west-1': 2}

Immutable configuration store — a tiny config object built entirely on tuples internally, guaranteeing nothing can silently rewrite a setting after construction:

class ImmutableConfig:
    def __init__(self, **kwargs):
        self._data = tuple(kwargs.items())

    def get(self, key):
        for k, v in self._data:
            if k == key:
                return v
        return None

>>> config = ImmutableConfig(host="localhost", port=8080)
>>> config.get("host"), config.get("port")
('localhost', 8080)

Quick Interview Answer

“These exercises combine the chapter’s core ideas into small, realistic code: tuple concatenation as the ‘append’ equivalent, the packing/unpacking swap idiom, and multi-value returns via an implicit tuple. The mini projects apply the same pattern to real infrastructure use cases — a server record unpacked into named fields for a description string, mapping a fixed region tuple to indices for round-robin logic, and an ImmutableConfig class that stores its settings as a tuple of (key, value) pairs internally specifically so nothing downstream can silently rewrite a setting after construction.”

Common Mistakes

  • Using t = t + (x,) repeatedly in a hot loop to “grow” a tuple — each concatenation allocates and copies the whole thing; if the collection genuinely needs to grow, use a list and convert to a tuple once at the end.
  • Forgetting the trailing comma when building a single-value tuple inside region_mapper-style code, silently producing the wrong type (see 11.2 Creating Tuples).
  • Implementing ImmutableConfig.get() with a linear scan over many settings when a dict would be both simpler and faster — a tuple of pairs communicates immutability, but doesn’t have to be the only internal representation if lookup performance matters.

Tuples: Chapter Practice

Mini lab

Unpack a host and port pair and print both values.

Expected output: localhost 8080

HintUse the chapter examples to choose the operation, then run your version before opening the solution.
Show solution
endpoint = ("localhost", 8080)
host, port = endpoint
print(host, port)

Knowledge check

Can a tuple contain a mutable list?

Check your answerYes. The tuple references cannot be replaced, but the referenced list can change.

Finish the chapter

  • Explain each operation without reading the lesson.
  • Change the sample input and predict the output before running it.
  • Revisit the chapter if the result differs from your prediction.

Back to this chapter · Choose the next chapter

12.1 Introduction to Sets

Python Sets

What Is a Set?

What Is It?

An unordered collection of unique, hashable values. Written with curly braces, or via set().

>>> unique_ports = {80, 443, 8080}
>>> unique_ports
{80, 443, 8080}

Sets were already introduced briefly in 5.6 Set Data Types; this chapter covers them in full.

Why Does It Matter?

Whenever “does this exist?” or “remove duplicates” matters more than order or position, a set is the right tool — this is a genuinely different job from what a list or tuple is built for.

Characteristics of Sets

Why Use Sets?

Two things a set does better than any other built-in type: automatic deduplication, and near-instant membership testing — O(1) average, versus a list’s O(n) (see 10.6 List Operators for the list-side comparison). Both are covered in depth in 12.13 Performance.

Real-World Applications

  • Deduplicating a list of email addresses
  • Tracking which users have already been notified
  • Finding common tags between two articles
  • Validating a value against an allowed set of options

Sets in DevOps

Deduplicating IP addresses from logs, comparing installed vs. required packages, diffing security group rules — 12.16 Sets in DevOps is dedicated entirely to these patterns.

>>> unique_ips = {"10.0.0.1", "10.0.0.2", "10.0.0.1"}
>>> unique_ips
{'10.0.0.1', '10.0.0.2'}

Quick Interview Answer

“A set is an unordered collection of unique, hashable values, built for exactly two jobs: automatic deduplication and fast membership testing. Internally it’s a hash table (see 12.3 Internal Representation), so checking x in s is O(1) average instead of the O(n) a list requires — that performance difference is the entire reason the type exists. The trade-offs that follow directly from being hash-based: no indexing, no guaranteed iteration order, and every element must itself be hashable.”

Common Mistakes

  • Reaching for a list and manually checking in repeatedly, when converting to a set once up front would make each check O(1) instead of O(n) (see 12.18 Best Practices).
  • Assuming a set preserves insertion order the way a list, tuple, or (since Python 3.7) a dict does — it doesn’t.
  • Writing s = {} expecting an empty set — that creates an empty dict instead (see 12.2 Creating Sets).

12.2 Creating Sets

Empty Set

Must use set() — {} creates an empty dict, not a set. This is 12.17 Common Mistakes’ most common gotcha.

>>> s = set()
>>> s
set()

Set Literals

>>> s = {1, 2, 3}

set() Constructor

Converts any iterable into a set, automatically deduplicating — the same conversion mechanics covered generally in 8.6 Collection Type Conversion.

>>> set([1, 1, 2, 3])
{1, 2, 3}

Sets from Lists

>>> set([1, 2, 2, 3])
{1, 2, 3}

Sets from Tuples

>>> set((1, 2, 3))
{1, 2, 3}

Sets from Strings

Produces a set of individual characters — same as list("...") does for lists (see 10.2 Creating Lists) — not a set of words.

>>> set("hello")
{'h', 'e', 'o', 'l'}

Quick Interview Answer

“The one rule that trips almost everyone up at least once: {} is a dict literal, not an empty set — curly braces were claimed by dict syntax first, so an empty set has to be written set() explicitly. Every other creation path — a literal {...}, or set() around a list, tuple, or string — automatically deduplicates. The string case deserves a specific callout: set("hello") doesn’t produce a set of words, it produces a set of individual characters, exactly the way list("hello") does.”

Common Mistakes

  • Writing s = {} and getting a dict back — always use set() for an empty set.
  • Expecting set("hello world") to split on whitespace into a set of words — it splits into individual characters instead; use set("hello world".split()) for words.
  • Forgetting set() on a list/tuple with duplicates silently drops them — useful when intended, a bug when the original count mattered.

12.3 Internal Representation

flowchart LR E["\"web01\""] -->|hash| H["hash value"] H -->|slot = hash % table_size| T["hash table (fixed-size slot array)"] T --> S0["slot 0"] T --> S1["slot 1 -> web01"] T --> S2["slot 2"] T --> S3["slot 3 -> db01"]

A set stores elements by hash, in a fixed-size slot array — not by insertion order.

Hash Table

Internally, a set is a hash table — each element’s position is computed from hash(element), not from insertion order. This is what makes membership testing O(1) average: checking for an item means computing its hash and looking at one slot, not scanning everything.

Memory Allocation

A set pre-allocates hash table space up front, even when empty — which is why an empty set and a small set can report similar memory overhead.

>>> import sys
>>> sys.getsizeof(set())
216
>>> sys.getsizeof({1, 2, 3})
216

References

Like lists and tuples, each set slot holds a reference to the actual object, not a copy of its data — the same reference model covered in 6.2 Objects and Variable References. The hash table itself stores hash values and pointers, not full copies of the elements.

Performance Benefits

The hash-table design is the entire reason sets exist as a distinct type from lists — full comparison in 12.13 Performance.

Quick Interview Answer

“A set is implemented as a hash table: every element’s slot is computed from hash(element) rather than from where it was inserted, which is exactly why membership testing is O(1) average — it’s a direct lookup, not a scan. That hash table is pre-allocated up front, which is why sys.getsizeof(set()) and a small non-empty set can report the same overhead. It also explains every other set behavior in this chapter: no indexing (there’s no meaningful position), no guaranteed iteration order (elements are laid out by hash, not insertion sequence), and a requirement that every element be hashable, since an unhashable element has no way to compute a slot at all.”

Common Mistakes

  • Assuming a set stores elements “in a list internally” — it’s a hash table, structurally closer to a dict’s key storage than to a list.
  • Expecting memory use to scale down for a near-empty set — the hash table’s base allocation is there whether the set holds 0 or 3 elements.
  • Forgetting the hash-table design is why set elements must be hashable — it’s not an arbitrary restriction, it’s a direct requirement of how the type is implemented.

12.4 Adding Elements

add()

Adds one element — silently does nothing if it’s already present (no error, no duplicate).

>>> s = {1, 2}
>>> s.add(3)
>>> s
{1, 2, 3}

update()

Adds every element from one or more other iterables — the set equivalent of 10.7 List Methods’ list.extend().

>>> s.update([4, 5], {6, 7})
>>> s
{1, 2, 3, 4, 5, 6, 7}

Adding Multiple Elements

update() accepts any number of iterables at once, of any iterable type (list, tuple, another set, …), merging them all in a single call.

Quick Interview Answer

“add() inserts exactly one element and is a silent no-op if it’s already present — there’s no error either way, which is a deliberate consequence of a set never having duplicates to begin with. update() is the multi-element equivalent, roughly mapping to list.extend(): it accepts any number of iterables in a single call and merges every element from each of them in, deduplicating automatically as it goes.”

Common Mistakes

  • Calling s.add([1, 2]) expecting it to add two elements — it tries to add the list [1, 2] as a single element and raises TypeError: unhashable type: 'list'; use update() for that instead.
  • Using add() in a loop to merge in another collection, when a single update() call does the same thing more directly.
  • Expecting add() on an existing value to raise an error or count occurrences — it’s always a silent no-op, since a set has no concept of duplicate membership.

12.5 Removing Elements

remove()

Removes a specific value — raises KeyError if it isn’t present.

>>> s = {1, 2, 3}
>>> s.remove(2)
>>> s
{1, 3}
>>> s.remove(99)
Traceback (most recent call last):
KeyError: 99

discard()

Same as remove(), but does not raise an error if the value is absent — the safer choice when it’s not certain the item exists.

>>> s.discard(99)      # no error, even though 99 isn't in the set
>>> s
{1, 3}

pop()

Removes and returns an arbitrary element — since sets are unordered (see 12.6 Set Indexing and Ordering), there’s no “first” or “last” to remove, unlike list.pop().

>>> p = s.pop()
>>> type(p)
<class 'int'>

clear()

>>> s2 = {1, 2, 3}
>>> s2.clear()
>>> s2
set()

del Statement

del doesn’t remove individual set elements (sets aren’t indexed) — it can only delete the entire set variable, same as for any other type.

>>> s3 = {1, 2, 3}
>>> del s3
>>> s3
Traceback (most recent call last):
NameError: name 's3' is not defined

Quick Interview Answer

“remove() and discard() do the identical removal — the only difference is that remove() raises KeyError if the value isn’t present, and discard() silently does nothing. That makes discard() the safer default whenever it isn’t already known the value exists. pop() is different from a list’s pop(): since a set has no order, there’s no ’last element’ to remove — it takes and returns an arbitrary one instead. And del never targets a set element the way it can a list index; it can only remove the variable binding itself, since sets aren’t indexed.”

Common Mistakes

  • Using remove() on a value that might not exist and getting an unhandled KeyError — use discard(), or check membership with in first.
  • Expecting pop() to remove a specific or “last-added” element — it removes whichever element the hash table happens to yield first, which is not predictable.
  • Trying del s[0] to remove an element — sets aren’t subscriptable at all (see 12.6 Set Indexing and Ordering); del only works on the variable name.

12.6 Set Indexing and Ordering

Why Sets Don’t Support Indexing

Sets have no concept of position — elements are located by hash, not by an integer offset, so s[0] is meaningless and raises TypeError.

>>> s = {3, 1, 2}
>>> s[0]
Traceback (most recent call last):
TypeError: 'set' object is not subscriptable

Iteration Order

The order elements appear when iterating a set is an implementation detail of the hash table layout, not insertion order — it can look consistent within one run but should never be relied upon.

>>> list({3, 1, 2})     # order here is incidental, not guaranteed by the language
[1, 2, 3]

Hashing Concept

Every element’s position in the underlying table is derived from hash(element) (see 12.3 Internal Representation) — this is exactly why set elements must be hashable, and why sets can’t contain lists or dicts.

Quick Interview Answer

“A set has no indexing because it has no positions — every element lives in a slot determined by its hash, not by where it was inserted, so s[0] raises TypeError: 'set' object is not subscriptable. The same reasoning explains iteration order: what you see when you loop over a set reflects the hash table’s internal layout, not insertion order, and while it can look stable within a single run, it is not something the language guarantees — sort explicitly with sorted(my_set) whenever the output order actually matters.”

Common Mistakes

  • Writing s[0] or slicing a set (s[:2]) out of list habit — neither works; convert to a sorted() list first if positional access is genuinely needed.
  • Assuming list(some_set) reproduces insertion order — it reflects hash-table layout instead, and can differ across Python versions or even between runs for certain types.
  • Relying on set iteration order in a script’s output, then being confused when it “changes” — it was never guaranteed to begin with.

12.7 Set Operators

Sets support real mathematical set operations directly as operators — the same conceptual operations from set theory, applied to Python collections.

| (Union)

Everything in either set (or both) — combines all unique elements.

>>> a, b = {1, 2, 3}, {2, 3, 4}
>>> a | b
{1, 2, 3, 4}

& (Intersection)

Only elements in both sets.

>>> a & b
{2, 3}

- (Difference)

Elements in the first set but not the second — order matters here, unlike union/intersection.

>>> a - b
{1}

^ (Symmetric Difference)

Elements in exactly one of the two sets (i.e. everything except the overlap).

>>> a ^ b
{1, 4}

Membership (in, not in)

The operation sets are famous for — O(1) average, versus O(n) for the same check on a list (see 10.6 List Operators).

>>> 2 in a, 5 not in a
(True, True)

Quick Interview Answer

“Sets support four operators straight out of set theory: | for union (everything in either), & for intersection (only what’s in both), - for difference (in the first but not the second — order matters here, unlike the other three), and ^ for symmetric difference (in exactly one, not both). Each has a method equivalent too (see 12.8 Set Methods), useful when chaining or passing as a callback. And in/not in membership testing is the operation sets exist for — O(1) average via a direct hash lookup, versus O(n) for the same check on a list.”

Common Mistakes

  • Assuming - is symmetric like | and & — a - b and b - a are generally different sets; only ^ is order-independent for “what differs.”
  • Forgetting that |, &, and - all build a new set rather than mutating either operand.
  • Reaching for a list and a manual loop to find common or missing elements between two collections, instead of converting to sets and using & or - directly (see 12.14 Common Algorithms).

12.8 Set Methods

Method equivalents of the operators in 12.7 Set Operators, plus relationship-testing methods that have no operator form.

union()

Same as | — method form is useful when chaining or passing as a function argument.

>>> {1, 2, 3}.union({2, 3, 4})
{1, 2, 3, 4}

intersection()

>>> {1, 2, 3}.intersection({2, 3, 4})
{2, 3}

difference()

>>> {1, 2, 3}.difference({2, 3, 4})
{1}

symmetric_difference()

>>> {1, 2, 3}.symmetric_difference({2, 3, 4})
{1, 4}

issubset()

True if every element of this set is also in the other set.

>>> {1, 2}.issubset({1, 2, 3})
True

issuperset()

True if this set contains every element of the other set — the inverse relationship of issubset().

>>> {1, 2, 3}.issuperset({1, 2})
True

isdisjoint()

True if the two sets share no elements at all.

>>> {1, 2}.isdisjoint({3, 4})
True

copy()

Returns a shallow copy — a new set object with the same elements.

>>> a = {1, 2, 3}
>>> c = a.copy()
>>> c is a, c == a
(False, True)

Quick Interview Answer

“Every operator from the previous section has a method twin — union(), intersection(), difference(), symmetric_difference() — useful mainly when chaining or passing as a function argument, where an operator can’t be used directly. The methods with no operator equivalent are the relationship tests: issubset() and issuperset() are inverses of each other, and isdisjoint() checks for zero overlap. copy() returns a genuine shallow copy — unlike a tuple’s copy.copy(), which can return the identical object since tuples are immutable, a set’s copy() always produces a distinct object, since a set is mutable and the copy must be safe to modify independently.”

Common Mistakes

  • Confusing issubset() and issuperset() — a.issubset(b) asks “is everything in a also in b?”; a.issuperset(b) asks the reverse.
  • Using a.difference(b) and expecting it to also remove those elements from a — it returns a new set; use difference_update() if in-place removal is actually wanted.
  • Assuming copy() performs a deep copy — like any shallow copy, if the set held mutable objects (which it normally can’t, since elements must be hashable) only the top-level container would be duplicated.

12.9 Built-in Functions

The same general-purpose functions used with lists (see 10.8 Built-in Functions) and tuples (see 11.7 Built-in Functions and Traversal) work identically on sets.

>>> len({3, 1, 4, 1, 5})
4
>>> max({3, 1, 4, 1, 5}), min({3, 1, 4, 1, 5})
(5, 1)
>>> sum({3, 1, 4, 1, 5})
13
>>> any({0, 0, 1}), all({1, 1, 1})
(True, True)

sorted()

Always returns a list — sets have no inherent order to sort in place.

>>> sorted({3, 1, 4, 1, 5})
[1, 3, 4, 5]

Quick Interview Answer

“Every general-purpose sequence function — len, max, min, sum, any, all — works on a set exactly as it does on a list or tuple, since none of them depend on order or indexing, only on iteration and comparison. sorted() is worth calling out specifically: it always returns a list, never a set, both because there’s no such thing as an in-place set.sort() and because the very concept of ‘sorted’ requires an order a set doesn’t have.”

Common Mistakes

  • Expecting sorted(some_set) to return a set — it always returns a list; wrap in set(sorted(...)) only if a set is genuinely needed afterward (which discards the ordering again).
  • Calling max()/min() on a set containing mixed, non-comparable types (like str and int together) — raises TypeError, same as it would for a list.
  • Forgetting sum() on a set of non-numeric hashable values (like strings) raises TypeError — sum() requires numeric elements, regardless of container type.

12.10 Traversing Sets

for Loop

>>> for x in {1, 2, 3}:
...     print(x, end=" ")
1 2 3

enumerate()

Works on a set like any iterable, but the indices it produces are meaningless as “positions” (see 12.6 Set Indexing and Ordering) — useful only for counting as you go, not for later lookup.

>>> for i, v in enumerate({10, 20, 30}):
...     print(i, v)
0 10
1 20
2 30

Iteration Best Practices

Never assume a set’s iteration order matches insertion order or any other meaningful sequence — if order matters for the output, convert to a sorted list first with sorted(my_set).

Quick Interview Answer

“Traversing a set looks identical to traversing a list — a plain for loop is the default, and enumerate() works exactly the same mechanically. The catch is entirely about meaning, not syntax: the index enumerate() produces for a set element is not a position it can be looked up by later, and isn’t tied to insertion order either — it’s only useful as a running counter. Whenever the traversal’s output actually needs a stable, meaningful order, convert to a sorted list first rather than relying on however the set happens to iterate.”

Common Mistakes

  • Using enumerate() over a set and later trying to access my_set[i] with that index — sets aren’t subscriptable at all.
  • Assuming two loops over the “same” set in one run will visit elements in the same order as each other — usually true within a single unmodified set, but not something to build logic around.
  • Printing or logging a set directly for output that other systems will parse, without sorting first — the unordered rendering can differ across runs or Python versions.

12.11 Frozen Sets

flowchart LR subgraph S["set {1, 2, 3}"] S1["Mutable"] S2["Not hashable"] end subgraph F["frozenset({1, 2, 3})"] F1["Immutable"] F2["Hashable"] end

Same operations as a regular set — union, intersection, membership — minus mutability.

What Is frozenset?

The immutable counterpart to set — same union/intersection/membership behavior, but no add(), remove(), or any method that would modify it. Because it’s immutable, it’s hashable, and therefore usable as a dict key or as a member of another set (which a regular set can never be) — the exact same trade-off covered for tuples in 11.9 Immutability and Copying.

Creating frozensets

>>> fs = frozenset([1, 2, 3])
>>> fs
frozenset({1, 2, 3})

Advantages

>>> fs.add(4)
Traceback (most recent call last):
AttributeError: 'frozenset' object has no attribute 'add'

>>> hash(fs)                    # hashable -- a regular set is NOT
-272375401224217160
>>> d = {fs: "value"}           # usable as a dict key
>>> d
{frozenset({1, 2, 3}): 'value'}

Use Cases

  • A fixed set of valid options that should never change at runtime
  • A set-of-sets, since a frozenset can be a member of another set
  • A dict key representing a group/combination of values

Quick Interview Answer

“frozenset is to set what a tuple is to a list: the same core behavior — union, intersection, membership — with mutation removed. That single change is what makes it hashable, and hashability is the whole payoff: a frozenset can be a dict key or a member of another set, both of which are flatly impossible for a regular set, since set.__hash__ doesn’t exist. The natural use case is a fixed group of values that itself needs to act like a value — a set-of-sets, or a dict key representing some combination.”

Common Mistakes

  • Trying fs.add(x) or fs.remove(x) on a frozenset — neither exists; AttributeError is raised immediately, the same shape of mistake as calling .append() on a tuple.
  • Reaching for a regular set as a dict key or set member and hitting TypeError: unhashable type: 'set' — swap in frozenset instead.
  • Assuming frozenset and tuple are interchangeable for “an immutable collection” — a frozenset still deduplicates and is unordered; a tuple preserves order and duplicates. Pick based on which behavior is actually needed.

12.12 Set Comprehensions

The same comprehension syntax as lists (see 10.9 Traversing and List Comprehensions), but with {} instead of [] — automatically deduplicates the result.

Basic

>>> {x**2 for x in range(5)}
{0, 1, 4, 9, 16}

Conditional

>>> {x for x in range(10) if x % 2 == 0}
{0, 2, 4, 6, 8}

Nested

A comprehension inside another expression — here, extracting a set of unique values from a nested structure.

>>> matrix = [[1, 2], [2, 3], [3, 4]]
>>> {x for row in matrix for x in row}
{1, 2, 3, 4}

Quick Interview Answer

“A set comprehension is a list comprehension with {} instead of [] — same for/if clauses, same nesting rules — with one automatic side effect: whatever the expression produces gets deduplicated as it’s built, since the result is a set. That makes it the natural one-liner for extracting unique values out of a nested or filtered structure, without a separate set(...) call wrapped around a list comprehension.”

Common Mistakes

  • Writing [x for x in ...] when a unique result is actually wanted, then wrapping the whole thing in set(...) afterward — a set comprehension does both steps in one pass.
  • Assuming a set comprehension preserves the order elements were generated in — like any set, the result’s iteration order isn’t guaranteed.
  • Using an unhashable expression result (like building a list per iteration) inside a set comprehension — raises TypeError: unhashable type, same as adding that value to a set directly.

12.13 Performance

flowchart TD A["x in big_list (10,000 items)"] --> A1["scans every element until found or exhausted"] A1 --> A2["O(n) -- 0.084s / 1000 checks"] B["x in big_set (10,000 items)"] --> B1["computes hash(x), checks one slot"] B1 --> B2["O(1) average -- 0.00005s / 1000 checks"]

Membership testing: a list scans linearly, a set jumps straight to the hash slot.

Time Complexity

OperationComplexityNotes
x in sO(1) averageDirect hash lookup
s.add(x)O(1) average
s.remove(x)O(1) average
s | t, s & t, s - tO(len(s) + len(t))Must scan both sets once
len(s)O(1)Cached, not recounted

Memory Usage

A set typically uses more memory per element than a list or tuple (see 12.3 Internal Representation) — the hash table needs extra space to keep lookups fast and collisions rare. This is the trade-off for O(1) membership testing.

Membership Testing

The headline performance case: for repeated “is this in my collection?” checks, converting to a set first pays for itself almost immediately on any non-trivial collection size.

>>> import timeit
>>> big_list = list(range(10000))
>>> big_set = set(big_list)
>>> timeit.timeit(lambda: 9999 in big_list, number=1000)
0.08375     # seconds -- illustrative, will vary by machine
>>> timeit.timeit(lambda: 9999 in big_set, number=1000)
0.00005

Quick Interview Answer

“Membership testing is the number one reason to pick a set: x in s is O(1) average because it’s a direct hash lookup, versus O(n) for a list, which has to scan element by element until it finds a match or runs out. That advantage compounds with every repeated check, which is why converting a list to a set once, up front, before doing thousands of in checks against it, is one of the most reliable performance wins in ordinary Python code. The trade-off is memory: a set’s hash table reserves extra space to keep collisions rare, so it typically costs more per element than an equivalent list or tuple — a worthwhile trade whenever membership testing happens more than a handful of times.”

Common Mistakes

  • Repeatedly checking x in my_list inside a loop against a list that never changes, instead of converting it to a set once beforehand (see 12.18 Best Practices).
  • Assuming set operations like | and & are also O(1) — they’re O(len(s) + len(t)), since both sets must be scanned once to compute the result.
  • Choosing a set purely for memory efficiency — for pure storage with no membership testing, a list or tuple is typically the more compact choice.

12.14 Common Algorithms

Remove Duplicates

The single most common use of a set — convert to a set and back to a list. If insertion order must be preserved instead, dict.fromkeys() is the standard alternative, since dicts have preserved insertion order since Python 3.7.

def remove_duplicates(lst):
    return list(set(lst))

>>> sorted(remove_duplicates([1, 2, 2, 3, 1]))
[1, 2, 3]

Unique Elements

def unique_elements(lst):
    return set(lst)

>>> unique_elements([1, 2, 2, 3])
{1, 2, 3}

Common Items

Finding overlap between two lists — convert both to sets and intersect.

>>> a, b = [1, 2, 3, 4], [3, 4, 5, 6]
>>> set(a) & set(b)
{3, 4}

Difference Between Lists

>>> set(a) - set(b)     # in a but not b
{1, 2}

Quick Interview Answer

“Four problems reduce to a one-liner once a list is converted to a set: deduplicating (list(set(lst)), or dict.fromkeys(lst) if order needs to be kept), finding unique elements (set(lst) directly), finding what two lists have in common (set(a) & set(b)), and finding what’s only in one (set(a) - set(b)). The pattern behind all four is the same — the moment ‘does this exist across both collections’ becomes the question, converting to sets and using the operators from 12.7 Set Operators is both shorter and asymptotically faster than nested loops.”

Common Mistakes

  • Deduplicating with list(set(lst)) when the original order matters — a set discards order entirely; use dict.fromkeys(lst) (then wrap in list(...) if a list is needed) to deduplicate while preserving first-seen order.
  • Solving “common items between two lists” with a nested loop (for x in a: for y in b: ...) — that’s O(n × m); set(a) & set(b) is O(n + m) and far more readable.
  • Forgetting set(a) - set(b) is directional — it answers “what’s in a but not b,” not the reverse; swap the operands for the other direction.

12.15 File Handling

Every line read from a file arrives as str (see 8.12 Type Conversion in File Handling) — a set comprehension is the standard one-line pattern for turning those lines directly into unique values.

Reading Data into Sets

>>> with open("ips.txt") as f:
...     unique_ips = {line.strip() for line in f}
...
>>> unique_ips
{'10.0.0.1', '10.0.0.2'}

Writing Sets

Since sets have no guaranteed order (see 12.6 Set Indexing and Ordering), sort before writing if the output order matters for the reader.

with open("unique_ips.txt", "w") as f:
    for ip in sorted(unique_ips):
        f.write(ip + "\n")

CSV Deduplication

Using a set comprehension while reading a CSV to drop duplicate rows/values in a single pass.

>>> import csv
>>> with open("data.csv") as f:
...     unique_values = {row[0] for row in csv.reader(f)}
...
>>> unique_values
{'a', 'b'}

Quick Interview Answer

“Reading unique values from a file is a one-line set comprehension over the lines — {line.strip() for line in f} — which both strips the trailing newline and deduplicates in the same pass. The one thing to remember when writing a set back out is that it has no order to preserve, so the output should be sorted explicitly (sorted(unique_ips)) rather than trusting iteration order to be stable or meaningful for whoever reads the file next. The same comprehension pattern works directly against csv.reader()’s rows for one-pass CSV deduplication.”

Common Mistakes

  • Writing a set’s contents to a file by iterating it directly without sorting — the row order in the output file can vary between runs.
  • Forgetting every value read from a text file or CSV is a str, even once collected into a set — numeric comparisons or arithmetic on the set’s contents need explicit conversion first.
  • Re-reading the same file into a fresh set every time inside a loop instead of building the set once outside the loop — repeated, unnecessary I/O.

12.16 Sets in DevOps

Five places sets are the natural fit in infrastructure code — specifically because deduplication and fast comparison matter, the same reason tuples show up in 11.12 Tuples in DevOps for fixed records and lists show up in 10.15 Lists in DevOps for ordered collections.

Unique IP Addresses

>>> log_ips = ["10.0.0.1", "10.0.0.2", "10.0.0.1", "10.0.0.3"]
>>> set(log_ips)
{'10.0.0.1', '10.0.0.2', '10.0.0.3'}

Installed Packages

Comparing what’s installed against what’s required with set difference — instantly reveals missing packages.

>>> installed = {"nginx", "redis", "curl"}
>>> required = {"nginx", "redis", "postgres"}
>>> required - installed     # missing packages
{'postgres'}

Security Groups

Finding ports allowed by both of two security groups via intersection — useful when auditing overlapping rule sets.

>>> sg1 = {"22", "80", "443"}
>>> sg2 = {"80", "443", "8080"}
>>> sg1 & sg2
{'80', '443'}

Duplicate Log Detection

Tracking a running “seen” set while scanning lines to flag exact repeats — a lightweight duplicate-log detector.

seen, duplicates = set(), set()
for line in ["a", "b", "a", "c"]:
    if line in seen:
        duplicates.add(line)
    seen.add(line)

>>> duplicates
{'a'}

Inventory Comparison

Comparing a current server fleet against the expected one — reporting what’s missing and what’s extra with two differences run in opposite directions.

>>> current = {"web01", "web02", "db01"}
>>> expected = {"web01", "web02", "web03"}
>>> print("missing:", expected - current)
missing: {'web03'}
>>> print("extra:", current - expected)
extra: {'db01'}

Quick Interview Answer

“Sets show up in DevOps tooling anywhere the question is ‘what’s unique,’ ‘what’s missing,’ or ‘what overlaps’ — deduplicating IPs pulled from a log file, diffing an installed-package set against a required one to find gaps, intersecting two security groups’ ports to audit overlapping access, and comparing a current server fleet against an expected inventory to report both what’s missing and what’s unexpectedly extra. The common thread is that each of these is a single set operation instead of nested loops: required - installed for missing packages, sg1 & sg2 for shared ports, two differences run in both directions for a full inventory diff.”

Common Mistakes

  • Using nested for loops to compare two lists of servers or packages, instead of converting both to sets and using -, &, or ^ directly.
  • Reporting only expected - current for an inventory comparison and forgetting current - expected — the first shows what’s missing, the second shows what’s unexpectedly extra; a complete audit usually needs both.
  • Deduplicating IP addresses or package names with a set, then needing the original order back for a report — sets discard order, so keep (or re-sort) a separate ordered structure if the sequence matters downstream.

12.17 Common Mistakes

Using {}

{} creates an empty dict, not an empty set — a legacy of dict literal syntax getting first claim on curly braces (see 12.2 Creating Sets). Always use set() for an empty set.

>>> empty = {}
>>> type(empty)
<class 'dict'>
>>> actual_empty_set = set()
>>> type(actual_empty_set)
<class 'set'>

Unhashable Types

Trying to put a mutable type (list, dict, another set) into a set fails, because set elements must be hashable (see 12.6 Set Indexing and Ordering) — use a tuple instead if a fixed, hashable grouping is needed.

>>> s = {[1, 2], 3}
Traceback (most recent call last):
TypeError: unhashable type: 'list'

Ordering Assumptions

Assuming list(some_set) will match the order elements were added — it won’t reliably. Sort explicitly whenever output order matters.

Quick Interview Answer

“Three mistakes account for most set bugs. First, {} — it’s a dict literal, not an empty set, purely because curly braces were claimed by dict syntax first; set() is required instead. Second, putting an unhashable value like a list or another (non-frozen) set into a set — raises TypeError: unhashable type immediately, because a set can’t compute a hash slot for something whose contents could still change. Third, and the subtlest, is assuming iteration order matches insertion order — it doesn’t, it reflects the hash table’s internal layout, and code that depends on it will eventually break in a way that’s hard to reproduce.”

Common Mistakes

  • Writing config = {} intending an empty set and later calling .add() on it, only to hit AttributeError because it’s actually a dict.
  • Putting a list of tags into a set of tags directly ({["prod", "web"]}) instead of a tuple ({("prod", "web")}) or individual hashable strings.
  • Building output that depends on set iteration order without sorting first, then debugging “inconsistent” results that were never actually guaranteed to be consistent.

12.18 Best Practices

When to Use Sets

  • Automatic deduplication is needed.
  • Fast, repeated membership testing is needed.
  • Mathematical set operations (union, intersection, difference) are needed.

Efficient Membership Tests

If a collection is checked with in more than a handful of times, convert it to a set once up front rather than testing against a list repeatedly (see 12.13 Performance).

# Less efficient -- O(n) per check, repeated many times
allowed = ["us-east-1", "us-west-2", "eu-west-1"]
for region in incoming_regions:
    if region in allowed:      # O(n) every single time
        ...

# More efficient -- convert once, O(1) per check afterward
allowed_set = set(allowed)
for region in incoming_regions:
    if region in allowed_set:  # O(1) average every time
        ...

Choosing Set vs. List vs. Frozenset

SituationBest choice
Order matterslist
Duplicates are meaningfullist
Need uniqueness + fast lookupset
Need to be a dict key / set memberfrozenset (not set)

Quick Interview Answer

“The decision to reach for a set comes down to two questions: does this need to be deduplicated automatically, and will it be checked with in more than once or twice? If either is true, a set is the right default — and the moment a collection is checked repeatedly against a fixed set of allowed values, converting it to a set exactly once up front is one of the cheapest performance wins available. The one caveat: if that immutable, fixed collection itself needs to be a dict key or live inside another set, frozenset is the correct choice, not a regular set, since a plain set can never be hashed.”

Common Mistakes

  • Testing membership against a list inside a loop that runs many times, instead of converting to a set once before the loop starts.
  • Reaching for a set when the order elements were added actually matters for later output — a set silently discards that information.
  • Using a regular set as a dict key or as a member of another set and hitting TypeError: unhashable type: 'set' — frozenset is the type built for exactly that use case (see 12.11 Frozen Sets).

12.19 Interview Questions

Conceptual Questions

  • What’s the difference between a set and a frozenset?
  • Why can’t a list be an element of a set?
  • Why is {} a dict and not a set?
  • What’s the time complexity of membership testing in a set, and why?
  • What’s the difference between remove() and discard()?

Scenario-Based Questions

  • Membership needs to be checked against a large collection thousands of times — what data structure would be used, and why?
  • A set of sets (a group of groups) is needed — why can’t regular sets be used as the inner elements, and what would be used instead?
  • Two lists represent “expected” and “actual” server inventories — how would the missing and extra entries be found, in one or two lines?

Coding Questions

  • Write a function that returns the unique elements common to three different lists.
  • Write a function that deduplicates a list while preserving original order (hint: sets alone won’t do this).
  • Given two sets of firewall rules, write code to report rules present in only one of them.
def common_to_three(a, b, c):
    return set(a) & set(b) & set(c)

>>> common_to_three([1, 2, 3], [2, 3, 4], [2, 3, 5])
{2, 3}

Quick Interview Answer

“Set questions cluster around three areas: the set-vs-frozenset distinction and why immutability buys hashability; why {} means dict and every set element must itself be hashable, which rules out lists and dicts as members; and the practical set-operator patterns — intersection for what’s common across collections, difference (run both directions) for what’s missing versus extra between an expected and actual inventory. The deduplicate-while-preserving-order question is the one that catches people off guard: a plain set discards order entirely, so dict.fromkeys(lst) — not set(lst) — is the answer whenever order needs to survive.”

Common Mistakes

  • Answering “sets and lists are basically the same, just pick either” without mentioning uniqueness, hashability requirements, or ordering.
  • Solving “common elements in three lists” with nested loops instead of chaining & across three sets in one expression.
  • Forgetting that set(lst) alone does not preserve order when asked to deduplicate a list while keeping the original sequence — dict.fromkeys(lst) is the actual answer.

12.20 Hands-on Exercises

Duplicate Remover

def remove_duplicates(items):
    return list(set(items))

>>> sorted(remove_duplicates([1, 1, 2, 3, 3]))
[1, 2, 3]

Unique Visitor Counter

Counting distinct visitors from a raw log of (possibly repeated) visitor IDs.

def unique_visitors(visitor_log):
    return len(set(visitor_log))

>>> unique_visitors(["u1", "u2", "u1", "u3"])
3

Package Comparator

Reporting missing and extra packages between an installed set and a required set — the same pattern as 12.16 Sets in DevOps.

def compare_packages(installed, required):
    installed_set, required_set = set(installed), set(required)
    return {
        "missing": required_set - installed_set,
        "extra": installed_set - required_set,
    }

>>> compare_packages(["nginx", "redis"], ["nginx", "postgres"])
{'missing': {'postgres'}, 'extra': {'redis'}}

Quick Interview Answer

“These three exercises build up from the simplest set use case to a realistic infrastructure pattern: remove_duplicates is list(set(...)) in its most basic form, unique_visitors shows that counting distinct items is just len() applied to a set instead of the raw log, and compare_packages combines two differences — required - installed and installed - required — into a single dict report, the exact shape a package-audit or inventory script needs in practice.”

Common Mistakes

  • Returning set(items) from remove_duplicates when a list was the expected return type — wrap with list(...) to match the original type.
  • Computing unique_visitors by manually looping and tracking a seen list with in checks, instead of the direct len(set(...)) one-liner.
  • Building compare_packages with only one direction of difference, silently missing either the “missing” or “extra” half of the comparison.

12.21 Mini Projects

Duplicate Log Analyzer

def analyze_duplicate_logs(lines):
    seen, dups = set(), set()
    for l in lines:
        if l in seen:
            dups.add(l)
        seen.add(l)
    return dups

>>> analyze_duplicate_logs(["a", "b", "a"])
{'a'}

Inventory Comparator

def compare_inventory(current, expected):
    c, e = set(current), set(expected)
    return {"missing": e - c, "extra": c - e}

>>> compare_inventory(["web01", "web02"], ["web01", "web03"])
{'missing': {'web03'}, 'extra': {'web02'}}

Firewall Rule Comparator

Symmetric difference is exactly the right tool for finding rules that differ between two rule sets, regardless of direction (see 12.7 Set Operators).

def compare_firewall_rules(rules_a, rules_b):
    return set(rules_a).symmetric_difference(set(rules_b))

>>> compare_firewall_rules(["22", "80"], ["80", "443"])
{'22', '443'}

Unique Host Tracker

def track_unique_hosts(connections):
    return set(connections)

>>> track_unique_hosts(["host1", "host2", "host1"])
{'host1', 'host2'}

Quick Interview Answer

“Each mini project applies one set operation to a realistic infrastructure task: a running seen set flags exact-repeat log lines as they’re scanned; two differences run in both directions turn a current and expected server list into a missing/extra report; symmetric_difference() finds every firewall rule that differs between two rule sets without needing to run the comparison twice in each direction; and a unique host tracker is set() doing its single most basic job — collapsing a connection log down to the distinct hosts involved.”

Common Mistakes

  • Implementing analyze_duplicate_logs by counting occurrences with a dict when only whether a line repeats is needed — the seen/dups two-set pattern is simpler and communicates intent more directly.
  • Reporting only expected - current (or only current - expected) from compare_inventory — a complete audit needs both directions.
  • Reaching for two separate differences (rules_a - rules_b and rules_b - rules_a) to find what differs between two rule sets, when a single symmetric_difference() call does the same thing in one step.

Sets: Chapter Practice

Mini lab

Find required tags absent from an observed tag set.

Expected output: [‘Owner’]

HintUse the chapter examples to choose the operation, then run your version before opening the solution.
Show solution
required = {"Name", "Owner", "Environment"}
observed = {"Name", "Environment"}
print(sorted(required - observed))

Knowledge check

Why sort the result for a report?

Check your answerSets do not guarantee iteration order. Sorting gives deterministic report output.

Finish the chapter

  • Explain each operation without reading the lesson.
  • Change the sample input and predict the output before running it.
  • Revisit the chapter if the result differs from your prediction.

Back to this chapter · Choose the next chapter

13.1 Dictionaries

Before you start

Work through the earlier fundamentals chapters. Use a Python 3 interpreter and run each example locally.

Core concepts

A dictionary maps unique, hashable keys to values. Assignment replaces an existing value for the same key. Dictionaries preserve insertion order; lookups are typically constant time on average.

Use direct indexing when a key is required and get() when a missing key is expected. A missing key accessed with brackets raises KeyError. Values can be mutable, so copying a dictionary does not recursively copy nested objects.

Worked example

server = {"name": "web-01", "state": "running"}
server["owner"] = "platform"
print(server.get("region", "unknown"))
for key, value in server.items():
    print(key, value)

Mini lab

Build a dictionary that counts occurrences in ["running", "stopped", "running"]. Expected result: {"running": 2, "stopped": 1}.

HintStart with the smallest input. Print intermediate values while exploring, then replace those prints with checks of the expected result.
Show a solution
counts = {}
for state in ["running", "stopped", "running"]:
    counts[state] = counts.get(state, 0) + 1
print(counts)

Knowledge check

What happens when you assign to an existing dictionary key?

Check your answerIts value is replaced; a second identical key is not added.

Common mistake

Never use a mutable list as a dictionary key. Use a tuple of hashable values if you need a composite key.

Completion checklist

  • Run the example and explain its output.
  • Complete the mini lab without copying the solution.
  • Test an edge case and explain how the code handles it.
  • Answer the knowledge check in your own words.

Fundamentals overview · Continue →

14.1 Conditionals

Before you start

Work through the earlier fundamentals chapters. Use a Python 3 interpreter and run each example locally.

Core concepts

An if branch runs when its condition is truthy. elif is checked only when earlier branches failed, and else covers the remaining cases. Indentation defines the body.

Empty collections, zero, None, and empty strings are falsy. Use is None when you specifically mean a missing value; a value of zero may still be valid. Compare values with ==, not is.

Worked example

cpu = 85
if cpu >= 90:
    print("critical")
elif cpu >= 70:
    print("warning")
else:
    print("healthy")

Mini lab

Classify usage values 0, 70, 90, and 100. Expected labels: healthy, warning, critical, critical. Reject values outside 0–100.

HintStart with the smallest input. Print intermediate values while exploring, then replace those prints with checks of the expected result.
Show a solution
for cpu in [0, 70, 90, 100]:
    if not 0 <= cpu <= 100:
        raise ValueError("CPU must be between 0 and 100")
    if cpu >= 90:
        print("critical")
    elif cpu >= 70:
        print("warning")
    else:
        print("healthy")

Knowledge check

Why should the >= 90 branch come before >= 70?

Check your answerA value of 95 also satisfies >= 70. The first matching branch wins.

Common mistake

Test values exactly on either side of a threshold. Avoid conditions that accidentally exclude zero.

Completion checklist

  • Run the example and explain its output.
  • Complete the mini lab without copying the solution.
  • Test an edge case and explain how the code handles it.
  • Answer the knowledge check in your own words.

Fundamentals overview · Continue →

15.1 Loops and Comprehensions

Before you start

Work through the earlier fundamentals chapters. Use a Python 3 interpreter and run each example locally.

Core concepts

A for loop visits an iterable’s items. range(n) yields integers from zero up to, but excluding, n. Use enumerate() when you need both an index and a value. A while loop repeats while a condition holds; make sure it can terminate.

break stops the loop and continue skips to the next iteration. A comprehension creates a new collection; prefer a normal loop when the body needs multiple steps or error handling.

Worked example

servers = ["web-01", "db-01", "web-02"]
web_servers = [name for name in servers if name.startswith("web-")]
for number, name in enumerate(web_servers, start=1):
    print(number, name)

Mini lab

From ports [22, 80, 443, 8080], collect only ports greater than 100. Expected: [443, 8080]. Repeat with an empty list.

HintStart with the smallest input. Print intermediate values while exploring, then replace those prints with checks of the expected result.
Show a solution
ports = [22, 80, 443, 8080]
selected = [port for port in ports if port > 100]
assert selected == [443, 8080]
assert [port for port in [] if port > 100] == []
print(selected)

Knowledge check

Does range(3) include the number 3?

Check your answerNo. It produces 0, 1, and 2.

Common mistake

Avoid removing items from a list while iterating over it. Build a filtered collection instead.

Completion checklist

  • Run the example and explain its output.
  • Complete the mini lab without copying the solution.
  • Test an edge case and explain how the code handles it.
  • Answer the knowledge check in your own words.

Fundamentals overview · Continue →

16.1 Functions

Before you start

Work through the earlier fundamentals chapters. Use a Python 3 interpreter and run each example locally.

Core concepts

Define a function with def, pass values through parameters, and use return to give a result to the caller. A function without an explicit return value returns None. Local names belong to that call unless you deliberately use a broader scope.

Keep filtering and validation separate from file or network operations so they can be tested easily. Type hints describe intended types but do not enforce them at runtime. Default arguments are evaluated once when the function is defined; avoid a mutable default such as items=[].

Worked example

def unhealthy(records, status="critical"):
    return [row for row in records if row.get("status") == status]

rows = [{"name": "web-01", "status": "critical"}]
print(unhealthy(rows))

Mini lab

Write is_healthy(status) returning a boolean. Check healthy, critical, and an empty string. The function must return a value rather than print it.

HintStart with the smallest input. Print intermediate values while exploring, then replace those prints with checks of the expected result.
Show a solution
def is_healthy(status):
    return status == "healthy"

assert is_healthy("healthy") is True
assert is_healthy("critical") is False
assert is_healthy("") is False

Knowledge check

Why use None instead of [] as the default for an optional list?

Check your answerA default list is shared across calls. Create a new list inside the function when the argument is None.

Common mistake

Do not mix printing and returning accidentally. Test the returned value directly.

Completion checklist

  • Run the example and explain its output.
  • Complete the mini lab without copying the solution.
  • Test an edge case and explain how the code handles it.
  • Answer the knowledge check in your own words.

Fundamentals overview · Continue →

17.1 Exceptions and Validation

Before you start

Work through the earlier fundamentals chapters. Use a Python 3 interpreter and run each example locally.

Core concepts

An exception interrupts normal execution. Catch the specific errors you can recover from with try and except. Use raise when input violates your function’s contract. Let unexpected failures propagate so they remain visible.

else runs if the try block succeeds; finally runs when leaving the block, including on failure. A context manager is usually the clearest way to close files. At a CLI boundary, translate known failures into a useful message and a nonzero exit code.

Worked example

def parse_port(text):
    port = int(text)
    if not 1 <= port <= 65535:
        raise ValueError("port must be between 1 and 65535")
    return port

try:
    print(parse_port("443"))
except ValueError as error:
    print(f"Invalid port: {error}")

Mini lab

Test parse_port with “443”, “abc”, and “70000”. The first returns 443; the others must raise ValueError.

HintStart with the smallest input. Print intermediate values while exploring, then replace those prints with checks of the expected result.
Show a solution
def parse_port(text):
    port = int(text)
    if not 1 <= port <= 65535:
        raise ValueError("invalid port")
    return port

assert parse_port("443") == 443
for text in ["abc", "70000"]:
    try:
        parse_port(text)
    except ValueError:
        pass
    else:
        raise AssertionError("Expected invalid port to fail")

Knowledge check

Should every exception be retried?

Check your answerNo. Invalid input and permission errors usually require a correction; retry only known transient failures with a bounded policy.

Common mistake

Avoid except Exception: pass. It can turn a failed operation into an apparent success.

Completion checklist

  • Run the example and explain its output.
  • Complete the mini lab without copying the solution.
  • Test an edge case and explain how the code handles it.
  • Answer the knowledge check in your own words.

Fundamentals overview · Continue →

18.1 Files, CSV, and JSON

Before you start

Work through the earlier fundamentals chapters. Use a Python 3 interpreter and run each example locally.

Core concepts

Use pathlib.Path for filesystem paths and with to close open files even when an error occurs. Specify UTF-8 when reading and writing text. Writing with mode w replaces an existing file; choose output paths deliberately.

Use csv.DictReader for header-based CSV records rather than splitting on commas: quoted fields may contain commas. CSV values are strings, so validate and convert fields. JSON supports nested objects and arrays; decoding untrusted or malformed input can fail.

Worked example

import csv
from io import StringIO

sample = "name,status\nweb-01,healthy\nweb-02,critical\n"
for row in csv.DictReader(StringIO(sample)):
    if row["status"] == "critical":
        print(row["name"])

Mini lab

Parse the sample CSV and serialize critical rows as JSON. Expected: [{"name": "web-02", "status": "critical"}]. Then try input with a header but no rows.

HintStart with the smallest input. Print intermediate values while exploring, then replace those prints with checks of the expected result.
Show a solution
import csv
import json
from io import StringIO

sample = "name,status\nweb-01,healthy\nweb-02,critical\n"
rows = [row for row in csv.DictReader(StringIO(sample))
        if row["status"] == "critical"]
print(json.dumps(rows))

Knowledge check

Why not parse CSV with line.split(",")?

Check your answerA quoted field may contain commas or newlines; a CSV parser handles the format correctly.

Common mistake

Validate required headers before processing records. Never overwrite the input file while you are reading it.

Completion checklist

  • Run the example and explain its output.
  • Complete the mini lab without copying the solution.
  • Test an edge case and explain how the code handles it.
  • Answer the knowledge check in your own words.

Fundamentals overview · Continue →

19.1 Modules and Imports

Before you start

Work through the earlier fundamentals chapters. Use a Python 3 interpreter and run each example locally.

Core concepts

A Python file can be imported as a module. Imported top-level code executes on first import in a process, so keep network calls and CLI parsing out of module scope. The if __name__ == "__main__" guard runs an entry point only when the file is executed directly.

Group related modules in a package, commonly using an __init__.py file. Prefer clear imports and avoid names such as csv.py or logging.py that shadow standard-library modules. Use a project virtual environment for third-party dependencies.

Worked example

# Save as health.py
def is_healthy(status):
    return status == "healthy"

def main():
    print(is_healthy("healthy"))

if __name__ == "__main__":
    main()

Mini lab

Save the example as health.py. In another file, import is_healthy and assert that critical returns False. Importing the module must not print anything.

HintStart with the smallest input. Print intermediate values while exploring, then replace those prints with checks of the expected result.
Show a solution
# Save beside health.py as check_health.py
from health import is_healthy

assert is_healthy("critical") is False
print("check passed")

Knowledge check

What does the main guard prevent?

Check your answerIt prevents the guarded entry-point code from running when another module imports the file.

Common mistake

Circular imports often indicate responsibilities are mixed. Move shared definitions to a small independent module.

Completion checklist

  • Run the example and explain its output.
  • Complete the mini lab without copying the solution.
  • Test an edge case and explain how the code handles it.
  • Answer the knowledge check in your own words.

Fundamentals overview · Continue →

20.1 Classes and OOP

Before you start

Work through the earlier fundamentals chapters. Use a Python 3 interpreter and run each example locally.

Core concepts

A class defines a type; an instance holds its own state. __init__ initializes an instance, and instance methods receive that instance as self. Prefer functions for simple transformations; use a class when related operations share meaningful state.

An instance attribute belongs to one object. A mutable class attribute is shared, which can surprise you. Composition builds an object from other objects; inheritance specializes a type and should preserve the parent’s expected behavior.

Worked example

class Server:
    def __init__(self, name, status):
        self.name = name
        self.status = status

    def is_healthy(self):
        return self.status == "healthy"

server = Server("web-01", "healthy")
print(server.is_healthy())

Mini lab

Create two Server objects. Change only the first status to critical. The second must remain healthy.

HintStart with the smallest input. Print intermediate values while exploring, then replace those prints with checks of the expected result.
Show a solution
class Server:
    def __init__(self, name, status):
        self.name = name
        self.status = status

    def is_healthy(self):
        return self.status == "healthy"

first = Server("web-01", "healthy")
second = Server("web-02", "healthy")
first.status = "critical"
assert first.is_healthy() is False
assert second.is_healthy() is True

Knowledge check

When would you prefer a function over a class?

Check your answerWhen an operation transforms input into output without maintaining shared state across calls.

Common mistake

Do not introduce an inheritance hierarchy just to reuse a few lines. Start with small functions and composition.

Completion checklist

  • Run the example and explain its output.
  • Complete the mini lab without copying the solution.
  • Test an edge case and explain how the code handles it.
  • Answer the knowledge check in your own words.

Fundamentals overview · Continue →