← World of AI
Artificial Intelligence
Computer Vision
Extracting meaning from images and video — detection, segmentation, tracking, recognition.
Vision moved from hand-designed features to learned ones with convolutional networks, and again toward transformer architectures that treat an image as a sequence of patches.
The tasks stack: classification says what is in the frame, detection says where, segmentation says which pixels, tracking says where it went next. Each step up demands more expensive labels, which is why self-supervised pre-training matters so much in this field.
Also in Artificial Intelligence
JOIN NOW
Begin the first module
It is free, it is the real curriculum, and if it is not for you, you have lost nothing but an evening.
Join any time · Build AI skills at your pace