Courseiva
Strings →hardMultiple Choice

PCAP Strings Practice Question

A script reads a binary file and decodes it as UTF-8. Some bytes are invalid UTF-8 sequences, causing a `UnicodeDecodeError`. The developer wants to replace invalid bytes with the replacement character U+FFFD. Which approach achieves this?

⚠ Common exam trap

Python Institute often tests the distinction between `errors='replace'` and `errors='ignore'`, where candidates mistakenly choose 'ignore' thinking it handles errors gracefully, but the trap is that 'ignore' silently drops invalid bytes instead of inserting a visible placeholder, which can lead to unintended data concatenation or loss of positional alignment.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

data.decode('utf-8', errors='replace')

The `errors='replace'` parameter in Python's `decode()` method replaces any bytes that cannot be decoded as valid UTF-8 with the Unicode replacement character U+FFFD, which is exactly what the developer wants. This approach ensures the script continues processing without raising a `UnicodeDecodeError` while preserving the overall structure of the data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    data.decode('utf-8', errors='strict')

    Why it's wrong here

    Using errors='strict' (the default) makes the decoder raise UnicodeDecodeError immediately upon encountering any byte sequence that is invalid UTF-8. Since the data came from a binary file, such invalid sequences are likely, and the exception would abort the script before any further processing. This mode does no recovery or substitution, so it fails to produce output rather than normalizing the problematic bytes.

  • ✗

    data.decode('utf-8', errors='surrogateescape')

    Why it's wrong here

    The surrogateescape error handler substitutes each invalid byte with a surrogate code point in U+DC80..U+DCFF, preserving the raw bytes for potential lossless re-encoding. However, these surrogates are not valid Unicode scalar values and will cause errors when the string is printed or encoded again with standard codecs. This is intentionally designed for OS filename round-tripping, not for producing clean readable text, so it does not achieve the intended replacement with U+FFFD.

  • ✓

    data.decode('utf-8', errors='replace')

    Why this is correct

    The replace handler maps each invalid byte or byte sequence to the Unicode replacement character U+FFFD, yielding a valid, displayable string that clearly marks where decoding errors occurred. This is the correct choice because the script likely needs to render the text content without crashing, while still indicating corrupted or non-UTF-8 portions. It sacrifices the original byte values but maintains the overall structure and length approximation of the data.

  • ✗

    data.decode('utf-8', errors='ignore')

    Why it's wrong here

    Using errors='ignore' simply discards invalid byte sequences without inserting any marker, so the resulting string silently loses chunks of the original data. This can cause adjacent valid characters to merge, potentially changing the meaning of the text. Unlike replace, ignore provides no visual indication that data was dropped, making it unsuitable when you need to preserve awareness of unclean input.

About these practice questions

One of 421 original PCAP practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PCAP practice question is part of Courseiva's free Python Institute certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCAP exam.