What is semantic segmentation in computer vision?
Semantic segmentation is a dense prediction task in which each pixel is assigned a class label from a fixed set of semantic categories, such as person, car, or background. Models like fully convolutional networks, U-Net, or DeepLab use spatial features and upsampling to produce an output map with the same resolution as the input. This pixel-wise classification gives a detailed scene understanding that is more granular than object-level or block-level annotations.
Why this answer
Semantic segmentation is a computer vision task that assigns a class label to every single pixel in an image, effectively partitioning the image into regions that correspond to different semantic categories (e.g., road, car, pedestrian). This is distinct from object detection, which only provides bounding boxes around objects, and from image captioning or OCR, which operate at a higher or different level of abstraction.
Exam trap
The trap here is that candidates often confuse semantic segmentation with object detection (Option A) because both involve identifying objects, but segmentation requires pixel-level precision rather than bounding boxes.
How to eliminate wrong answers
Option A is wrong because detecting boundaries of objects using rectangular boxes describes object detection, not semantic segmentation, which operates at the pixel level rather than with bounding boxes. Option C is wrong because generating natural language descriptions of images is image captioning, a different computer vision task that produces text, not pixel-level classification. Option D is wrong because extracting text from images using OCR is optical character recognition, which focuses on text extraction, not pixel-wise semantic labeling.