If you cannot quickly search through a 10,000-line log file to find a single error, you will spend hours scrolling and missing the clue that crashed a company's server. Text processing with filters and regular expressions is the skill that turns messy, overwhelming output into precise, actionable information. For the LPIC-1 exam, mastering these tools — grep, sed, and awk — is essential because they appear in a full third of the system administration tasks you will be tested on.
Jump to a section
A simple way to picture Text Processing with Filters and Regular Expressions
Picture a crowded beach after a bank holiday weekend: sand strewn with bottle caps, lost keys, a dropped phone, and a single gold ring buried three inches deep.
The beachcomber's metal detector is the filter: a tool that cuts through the noise. You sweep it over the sand, and it beeps only when it senses certain metals. That beep is a match. Now imagine you don't just want any metal — you want only rings, not foil or coins. You tune the detector's sensitivity and discrimination settings. Those settings are your regular expression: a pattern that tells the detector, "Sound off only for objects that are gold-coloured, loop-shaped, and about 2-3cm wide."
When the detector beeps, you dig down and find the ring. But what if you need to report how many rings you found today? You pull out a waterproof phone, scan the beach's finds into a list, then filter that list down to rings, count them, and write a report. In the IT version, the detector is grep, the discrimination pattern is a regex, and the report loop is awk. The beach is your log file — a chaotic text stream that holds treasure if you know how to search it.
Text processing is the art of manipulating streams of characters — lines of text — that flow between commands, into files, or out of programs. On a Linux system, almost everything is text: configuration files, logs, command output, user data. Filters are programs that read text input, transform it in some way, and then write the result to output, often of their own accord without needing a separate output command. A classic filter is a pipe filter: you connect the output of one command into the input of another using the pipe symbol '|'. For example, 'ls -l | sort' sends the detailed file listing into the sort program, which organises the lines alphabetically.
The three key filter tools for the LPIC-1 exam are grep, sed, and awk.
grep (global regular expression print) searches for lines that contain a match to a specified pattern. It is like the find function in a document, but for entire files and across many files at once. By default, grep prints the whole line that matched. You can use options like '-i' for case-insensitive search, '-v' to invert the match (show lines that do NOT match), and '-c' to count matching lines instead of printing them. The power of grep comes from its use of regular expressions, which are patterns written in a special language.
Regular expressions (regex or regexp) are a notation for describing text patterns. A basic regex uses literal characters — the letter 'a' matches the letter 'a' — but also metacharacters that have special meanings. The dot '.' matches any single character. The asterisk '*' means 'zero or more of the preceding character'. The caret '^' anchors the pattern to the start of a line, and the dollar sign '$' anchors it to the end. Square brackets '[abc]' define a character class, matching any one of the characters inside. For example, the pattern '^error.*[0-9]' would match any line that starts with 'error', followed by any characters, ending with a digit.
sed (stream editor) reads a stream of text line by line and performs editing operations on it. You can think of sed as a find-and-replace tool on steroids. Its most common use is substitution: 's/old/new/' replaces the first occurrence of 'old' with 'new' on each line. With the 'g' flag at the end — 's/old/new/g' — it replaces every occurrence on each line. Sed can also delete lines that match a pattern with the 'd' command, insert or append lines, and even run multiple commands with the '-e' option. Unlike a text editor like nano, sed works entirely from the command line and is ideal for batch processing scripts.
awk is a full programming language designed for text processing, but you can use it simply as a filter. The name comes from its creators: Aho, Weinberger, and Kernighan. Awk works on a record-and-field model: each line is a record, and each word on that line (separated by whitespace by default) is a field. You refer to fields with $1, $2, $3 — $0 is the whole line. For example, 'awk '{print $1, $3}'' prints only the first and third fields of each line. Awk can also evaluate conditions: 'awk '$3 > 50 {print $1}'' prints the first field of lines where the third field is greater than 50. It supports if statements, loops, and built-in variables like NR (line number) and NF (number of fields).
Why do these tools exist? Before graphical interfaces, everything was text. System administrators needed ways to extract specific pieces of information from large texts. Grep, sed, and awk are the descendants of the Unix philosophy of small, composable tools that do one thing well. They replaced manual scanning by humans, which was slow and error-prone. Today, they are still the fastest way to search, edit, and summarise text on a server — especially when you cannot use a GUI because the server is headless (no monitor) or you are logged in over SSH.
To combine these tools, you use pipes. For instance, pipe your logs into grep to extract error lines, pipe those into sed to clean up timestamps, then pipe into awk to reformat the output. This chain is a pipeline, and each stage is a filter. Understanding this flow is the core of the LPIC-1 objective 103.2.
Identify the Input Source
Decide where your text data is coming from: a file (like /var/log/syslog), the output of another command (like 'ps aux'), or direct user input. This determines how you start your pipeline. For example, 'cat file.txt' or 'command' then pipe.
Select the Right Filter Tool
Determine what operation you need: searching (grep), editing (sed), or extracting fields (awk). If you need just a line that contains a keyword, grep is simplest. If you need to replace 'old' with 'new', use sed. If you need to sum a column of numbers, use awk.
Write the Pattern or Expression
Craft your regex or action. Start simple and test with a small sample. For grep, 'grep 'error''. For sed, 's/error/warning/'. For awk, '{print $2}'. Use single quotes around the pattern to prevent shell expansion.
Test in Isolation on a Subset
Run the command on a small piece of data first, maybe using 'head' to grab the first 10 lines: 'head -10 file.txt | grep 'pattern''. This lets you verify the pattern matches as expected without waiting for large file processing.
Build and Execute the Pipeline
Chain filters using pipes. For instance: 'grep -i 'error' log.txt | sed 's/\[WARN\]//g' | awk '{print $1, $3}' | sort | uniq -c'. Each stage reduces or transforms the data. Test each stage incrementally, adding one pipe at a time.
Redirect or Record Output
Decide where the final output goes. If you want a file, use '> output.txt'. If you want to see it on screen, let it flow to stdout. For in-place editing with sed, add the -i option with care. Always review the output to confirm correctness.
An IT professional manages a web server that runs an e-commerce site. One morning, customers report that the checkout page is returning a 500 Internal Server Error. The admin needs to find the cause fast.
She SSHes into the server and opens the application log file, app.log, located in /var/log/. The file is 200 megabytes — far too large to read manually. She starts by searching for lines related to the error. She runs:
grep '500' /var/log/app.log
This returns dozens of lines. Too many. She needs to narrow it down. The pattern '500' could appear in many contexts — status codes, error messages, even memory addresses. So she refines the search to look for the phrase '500 Internal Server Error' and, more specifically, restricts it to errors from the payment module:
grep 'payment.*500' /var/log/app.log
That regex matches lines containing 'payment' followed by any characters, then '500'. She gets five lines. Each line has a timestamp, a severity level, and a description. However, the timestamps are in a format like '2025-03-15 14:32:01', and the descriptions have extra debug symbols that make the output cluttered. She pipes the result through sed to strip the debug symbols:
grep 'payment.*500' /var/log/app.log | sed 's/\[DEBUG\]//g'
Now each line is cleaner, but she wants to see only the timestamp and the error message, ignoring fields like process ID and thread number. She pipes the output into awk to print only those columns:
grep 'payment.*500' /var/log/app.log | sed 's/\[DEBUG\]//g' | awk '{print $1, $2, $7}'
The output shows the date, time, and the seventh field, which happens to be the error description. She spots a pattern: the error only occurs when a specific user ID appears. She can then extract lines where the user ID field equals a certain value by adding an awk condition: 'awk '$7 == "user123"''.
Step by step, she has filtered a massive log file down to a single issue: the payment gateway's API returned a timeout for specific accounts. She reports this to the development team with the exact timestamp and affected user IDs. Her quick pipeline saved hours of manual inspection. In an exam scenario, you might be asked to construct such a pipeline or explain what each part does. The real power is in knowing which tool to use when: grep for searching, sed for transforming, awk for extracting and summarising.
The LPIC-1 exam tests objective 103.2 with multiple-choice questions, fill-in-the-blank, and simulation-based items. Expect questions that require you to: identify correct syntax for grep, sed, and awk; understand regex metacharacters and their meanings; distinguish between basic and extended regular expressions; and know when to use each tool.
Key concepts they love to test:
Basic vs extended regular expressions: grep by default uses basic regex (BRE), where metacharacters like '+', '?', '{', '|', '(', ')' must be escaped with a backslash to activate their special meaning. Use 'grep -E' for extended regex (ERE), where those characters are special without escaping. The exam will give you a pattern and ask which grep invocation correctly matches it.
The bracket expression: '[a-z]' matches any lowercase letter, but '[^a-z]' matches anything that is NOT a lowercase letter. The caret inside brackets is negation. Outside brackets, it anchors to the start of a line.
Sed substitution flags: the 'g' flag replaces all occurrences; without it, only the first per line. '2' replaces the second occurrence. 'p' prints the line after substitution (often used with -n). A common trap: sed does not edit the file in place by default — use -i for in-place editing, but the exam expects you to know that output goes to stdout unless redirected.
Awk pattern-action structure: 'pattern { action }'. If no pattern is given, the action runs on every line. If no action is given, the line is printed. A common question: what does 'awk 'length($0) > 80 {print $0}'' do? It prints lines longer than 80 characters.
The difference between filters and commands: A filter reads from stdin and writes to stdout. Commands like cat, head, tail, and sort are also filters, but grep, sed, awk are specialised.
Trap patterns to watch for: The exam may present a regex that looks correct but uses the wrong escaping. For example, the pattern 'hello|world' in basic grep (without -E) will match the literal string 'hello|world', not 'hello' OR 'world'. To match alternation, you need 'hello\|world' or use egrep/grep -E. Another trap: confusing grep options. '-i' is case-insensitive; '-v' inverts match; '-c' counts; '-o' prints only the matching part, not the whole line. They may ask: 'which option will show only the matched text?' Answer: -o.
Finally, practise with actual log files. The exam will present snippets and ask you to determine the output of a given command. The best way to prepare is to run these commands on real data and see the results. Memorise the syntax for sed substitutions: 's/pattern/replacement/flags'. Know that awk uses $1, $2, etc., and that $0 is the entire record. Understand that the field separator (FS) defaults to whitespace but can be set with -F or the FS variable. These are the core details that exam questions hammer.
Grep searches for lines matching a pattern and prints them, but never modifies the input file without redirection.
A regular expression is a pattern-matching language where metacharacters like ., *, ^, $, [ ] have special meanings.
Sed is a stream editor that applies editing commands line by line, most commonly the substitution command s/old/new/flags.
Awk works on a record (line) and field (word) model, using $1, $2, etc. to reference fields, and supports arithmetic and conditional logic.
In basic grep (without -E or egrep), metacharacters +, ?, {, |, (, ) must be escaped with a backslash to activate special meaning.
Pipes (|) connect the stdout of one filter to the stdin of the next, allowing you to combine grep, sed, and awk into powerful pipelines.
The -v option in grep inverts the match, showing only lines that do NOT contain the pattern.
Sed's -i option modifies the file in place, but always test without -i first using stdout to avoid accidental data loss.
These come up on the exam all the time. Here's how to tell them apart.
grep
Primarily used for searching and selecting lines that match a pattern
Outputs matched lines unchanged by default
Cannot modify text within a line without piping to another tool
sed
Primarily used for editing and transforming text line by line
Always applies an action (substitution, deletion, insertion)
Can change content within a line using substitute 's' command
sed
Designed for editing operations on lines
Uses simple commands like s, d, i, a
Lacks built-in arithmetic and field extraction by default
awk
Designed for data extraction and report generation
Full programming language with variables, arrays, and loops
Has built-in field splitting ($1, $2) and arithmetic
Basic Regex (BRE) in grep
Metacharacters like +, ?, |, (, ) are literal unless escaped with backslash
Backreferences like \1 work as usual
Default mode for grep without options
Extended Regex (ERE) in grep -E
Metacharacters +, ?, |, (, ) are special without escaping
Backreferences still work, but alternation is easier: 'cat|dog'
Activated with -E or using egrep command
awk default field separator (FS)
Whitespace (spaces and tabs) splits fields
Multiple whitespace characters treated as a single separator
Each word is a separate field
awk with custom field separator (-F)
Single character (e.g., -F: uses colon) splits fields
Colon is treated literally, sequence matters
Often used for /etc/passwd or CSV files
Mistake
Regular expressions are unique to grep and you only need to learn one set of characters.
Correct
Regex patterns exist in many tools (grep, sed, awk, vim, Python, Perl) with slight variations. The core syntax is similar, but flags and escaping differ. In grep, basic and extended regex modes behave differently. In awk, the regex is between slashes, and in sed, it appears in the substitution command.
Beginners often use grep's basic regex and then try the same pattern in awk or sed without adjusting escaping, leading to no matches.
Mistake
Sed always modifies the file you point it at.
Correct
By default, sed writes to stdout and does not touch the original file. Only the -i option (in-place) modifies the file directly. Without -i, you must redirect output to overwrite if needed.
Users who come from word processors assume editing commands alter the file. Sed's stream-oriented nature is counterintuitive to that model.
Mistake
Awk is just a filter like grep and can only print or not print lines based on a condition.
Correct
Awk is a full programming language with variables, arrays, loops, and arithmetic. It can compute sums, averages, generate reports, and even produce formatted output using printf. While it is often used as a simple filter, its programming capabilities are what make it powerful.
Because simple uses of awk mimic grep in functionality, beginners underestimate its depth and fail to use its more advanced features like built-in functions.
Mistake
The dot '.' metacharacter in regex matches any character including a newline.
Correct
The dot matches any single character except the newline character. To match across multiple lines, you need different constructs or tools.
Beginners test the dot on a string with a newline in the middle and get no match, then think regex is broken. This is a classic point of confusion.
Mistake
You can only use one filter at a time. Combining them is confusing and rarely needed.
Correct
Chaining filters with pipes is a core Unix principle. Almost all real-world text processing uses multiple tools together. For example, grep -> sed -> awk is a common pipeline.
New users often think each tool is standalone and that you must choose one, not realising that composing tools is where the real efficiency comes from.
Reveal each answer, then mark whether you got it right. Score 60%+ to unlock the next chapter.
grep uses basic regular expressions by default. egrep (or grep -E) uses extended regular expressions, which treat +, ?, |, (, ) as special without needing backslashes. fgrep (or grep -F) treats the pattern as a fixed string, ignoring all regex metacharacters — useful when searching for literal punctuation.
Use the -v flag with grep. For example: 'grep -v 'error' file.txt' will show all lines that do NOT contain the word 'error'.
Yes. Print the built-in NR (record number) variable: 'awk '{print NR, $0}' file.txt' will prefix each line with its line number.
Without 'g', sed replaces only the first occurrence of the pattern on each line. With 'g' (global), it replaces every occurrence on each line. Example: 's/old/new/g' replaces all instances of 'old' with 'new' per line.
By default, grep, sed, and awk work line by line. A pattern like 'start.*end' will not match if 'start' is on one line and 'end' on another. You need to use tools or options that handle multiline patterns, such as 'grep -z' (treats input as null-separated) or use a different approach like 'sed' with the N command.
Use 'grep -c 'pattern' file.txt' to count matching lines. To count total occurrences (including multiple per line), use 'grep -o 'pattern' file.txt | wc -l'.
You've finished Text Processing with Filters and Regular Expressions. Continue through the LPIC-1 study guide to build a complete picture of the exam.
Done with this chapter?