Recent from talks
Jitendra Malik
Knowledge base stats:
Talk channels stats:
Members stats:
Jitendra Malik
Jitendra Malik (born 11 October 1960) is an Indian-American academic who is the Arthur J. Chick Professor of Electrical Engineering and Computer Sciences at the University of California, Berkeley. He is known for his research in computer vision.
Malik was born in Mathura, India, on October 11, 1960. He did his schooling from Jabalpur, at the St. Aloysius Senior Secondary School. He received the BTech degree in electrical engineering from Indian Institute of Technology Kanpur in 1980 and the PhD degree in computer science from Stanford University in 1985. In January 1986, he joined the University of California, Berkeley, where he is currently the Arthur J. Chick Professor in the Computer Science Division, Department of Electrical Engineering and Computer Sciences (EECS). He is also on the faculty of the department of Bioengineering, and the Cognitive Science and Vision Science groups. He served as the chair of the Computer Science Division during 2002–2004 and as the department chair of EECS during 2004–2006 and 2016–2017.
He served as a visiting research scientist at Google during 2015–2016 and later joined Meta's Fundamental AI Research (FAIR), serving as Research Director and Site Lead in Menlo Park before becoming Vice President for Robotics Research in 2025. In 2026 he joined Amazon as Vice President and Distinguished Scientist at its Frontier AI and Robotics (FAR) laboratory while remaining affiliated with UC Berkeley.
Malik has supervised more than eighty doctoral students and postdoctoral researchers, many of whom have become leading academics and industrial researchers at institutions including MIT, UC Berkeley, Carnegie Mellon University, Cornell University, the University of Illinois Urbana–Champaign, the University of Pennsylvania, the University of Michigan, Google, Meta, and other major research organizations.
Malik has been a leader of computer vision across multiple decades, steering the field and contributing to many of its fundamental results. Malik’s approach to computer vision commonly draws inspiration and insight from psychology and neuroscience, and contributes to computational modeling of biological vision, for example his framing of the central problems of computer vision as the 3R’s, recognition, reconstruction and re-organization, emphasizing the close coupling between these. He was awarded the 2019 IEEE Computer Society’s Computer Pioneer Award for his “leading role in developing Computer Vision into a thriving discipline through pioneering research, leadership, and mentorship”. More recently, Malik’s group turned its attention to robotics, making significant contributions to navigation and legged locomotion.
The late 1990s marked a transition from geometric to learning techniques for visual recognition in computer vision. At this time, the techniques were from the statistical machine learning tradition – nearest neighbor, random forests and support vector machines. To apply these to visual data, handdesigned features were necessary and the Malik group pioneered features such as textons (vector quantized filter outputs) , shape contexts (relative arrangements of points) as well as novel machine learning techniques such as fast intersection kernels, SVM-KNN etc. This enabled them to achieve world record performance numbers for various tasks - handwritten digit recognition (2001), breaking CAPTCHAs (2002), Caltech101 categories (2005-07), people detection (2009). The shape context method [8] received the Helmholtz test-of-time award, and has more than 9000 citations. In 2012 a major paradigm shift occurred with the “AlexNet” work from Geoff Hinton’s group launching the deep learning revolution.
Malik played a catalytic role in this by encouraging Hinton to prove that deep learning worked better by competing on standard visual recognition benchmarks, specifically ImageNet. However, even after AlexNet, the broader computer vision community was still skeptical about the generality of the approach, since the ImageNet challenge was for classification and did not require object localization, for which at the time, PASCAL VOC was the accepted benchmark. Girshick et al invented the R-CNN method which proved that indeed this could be done, with a stage of pre-training on ImageNet classification followed by fine-tuning for object detection. This paper has more than 44000 citations for its CVPR 2014 version and received the Longuet-Higgins test-of-time award. A subsequent paper defined the problem of object instance segmentation and presented a model for its solution. R-CNN variants dominated object recognition for the better part of a decade until finally being superseded by VLMs.
Visual grouping and figure-ground discrimination were first studied by the Gestalt school of visual perception more than a century ago. The key insight, translated into modern terminology, is that we do not perceive an image as just a set of pixels, rather organize it into “regions” or “segments” corresponding to objects in the world. Operationalizing this computationally had been a central problem in computer vision from its early days. Malik’s group, over a period of two decades, 1995- 2015, made fundamental contributions which transformed the area from a miscellaneous collection of techniques to an empirically based science, and with the latest tools from deep learning, now largely a solved problem (cf. the Segment Anything system from Meta). Before Malik’s group’s work, contour detection and image segmentation research in computer vision was evaluated in an ad hoc manner. Authors showed their algorithms results without any measure of what is the right answer, unlike the case for mathematically well posed problems like structure from motion. Malik’s group changed the culture. They created the “Berkeley Segmentation Data Set” using multiple human observers to mark the boundaries they perceived in the image, and showing that the observers were consistent with a hierarchical model of segmentation. The creation of this dataset itself was noteworthy-using human observers to annotate images had previously been done for very niche categories like digits and faces, whereas here natural images were being segmented at scale. The BSDS paper has more than 10,000 citations and received the Helmholtz test-of-time award.
Hub AI
Jitendra Malik AI simulator
(@Jitendra Malik_simulator)
Jitendra Malik
Jitendra Malik (born 11 October 1960) is an Indian-American academic who is the Arthur J. Chick Professor of Electrical Engineering and Computer Sciences at the University of California, Berkeley. He is known for his research in computer vision.
Malik was born in Mathura, India, on October 11, 1960. He did his schooling from Jabalpur, at the St. Aloysius Senior Secondary School. He received the BTech degree in electrical engineering from Indian Institute of Technology Kanpur in 1980 and the PhD degree in computer science from Stanford University in 1985. In January 1986, he joined the University of California, Berkeley, where he is currently the Arthur J. Chick Professor in the Computer Science Division, Department of Electrical Engineering and Computer Sciences (EECS). He is also on the faculty of the department of Bioengineering, and the Cognitive Science and Vision Science groups. He served as the chair of the Computer Science Division during 2002–2004 and as the department chair of EECS during 2004–2006 and 2016–2017.
He served as a visiting research scientist at Google during 2015–2016 and later joined Meta's Fundamental AI Research (FAIR), serving as Research Director and Site Lead in Menlo Park before becoming Vice President for Robotics Research in 2025. In 2026 he joined Amazon as Vice President and Distinguished Scientist at its Frontier AI and Robotics (FAR) laboratory while remaining affiliated with UC Berkeley.
Malik has supervised more than eighty doctoral students and postdoctoral researchers, many of whom have become leading academics and industrial researchers at institutions including MIT, UC Berkeley, Carnegie Mellon University, Cornell University, the University of Illinois Urbana–Champaign, the University of Pennsylvania, the University of Michigan, Google, Meta, and other major research organizations.
Malik has been a leader of computer vision across multiple decades, steering the field and contributing to many of its fundamental results. Malik’s approach to computer vision commonly draws inspiration and insight from psychology and neuroscience, and contributes to computational modeling of biological vision, for example his framing of the central problems of computer vision as the 3R’s, recognition, reconstruction and re-organization, emphasizing the close coupling between these. He was awarded the 2019 IEEE Computer Society’s Computer Pioneer Award for his “leading role in developing Computer Vision into a thriving discipline through pioneering research, leadership, and mentorship”. More recently, Malik’s group turned its attention to robotics, making significant contributions to navigation and legged locomotion.
The late 1990s marked a transition from geometric to learning techniques for visual recognition in computer vision. At this time, the techniques were from the statistical machine learning tradition – nearest neighbor, random forests and support vector machines. To apply these to visual data, handdesigned features were necessary and the Malik group pioneered features such as textons (vector quantized filter outputs) , shape contexts (relative arrangements of points) as well as novel machine learning techniques such as fast intersection kernels, SVM-KNN etc. This enabled them to achieve world record performance numbers for various tasks - handwritten digit recognition (2001), breaking CAPTCHAs (2002), Caltech101 categories (2005-07), people detection (2009). The shape context method [8] received the Helmholtz test-of-time award, and has more than 9000 citations. In 2012 a major paradigm shift occurred with the “AlexNet” work from Geoff Hinton’s group launching the deep learning revolution.
Malik played a catalytic role in this by encouraging Hinton to prove that deep learning worked better by competing on standard visual recognition benchmarks, specifically ImageNet. However, even after AlexNet, the broader computer vision community was still skeptical about the generality of the approach, since the ImageNet challenge was for classification and did not require object localization, for which at the time, PASCAL VOC was the accepted benchmark. Girshick et al invented the R-CNN method which proved that indeed this could be done, with a stage of pre-training on ImageNet classification followed by fine-tuning for object detection. This paper has more than 44000 citations for its CVPR 2014 version and received the Longuet-Higgins test-of-time award. A subsequent paper defined the problem of object instance segmentation and presented a model for its solution. R-CNN variants dominated object recognition for the better part of a decade until finally being superseded by VLMs.
Visual grouping and figure-ground discrimination were first studied by the Gestalt school of visual perception more than a century ago. The key insight, translated into modern terminology, is that we do not perceive an image as just a set of pixels, rather organize it into “regions” or “segments” corresponding to objects in the world. Operationalizing this computationally had been a central problem in computer vision from its early days. Malik’s group, over a period of two decades, 1995- 2015, made fundamental contributions which transformed the area from a miscellaneous collection of techniques to an empirically based science, and with the latest tools from deep learning, now largely a solved problem (cf. the Segment Anything system from Meta). Before Malik’s group’s work, contour detection and image segmentation research in computer vision was evaluated in an ad hoc manner. Authors showed their algorithms results without any measure of what is the right answer, unlike the case for mathematically well posed problems like structure from motion. Malik’s group changed the culture. They created the “Berkeley Segmentation Data Set” using multiple human observers to mark the boundaries they perceived in the image, and showing that the observers were consistent with a hierarchical model of segmentation. The creation of this dataset itself was noteworthy-using human observers to annotate images had previously been done for very niche categories like digits and faces, whereas here natural images were being segmented at scale. The BSDS paper has more than 10,000 citations and received the Helmholtz test-of-time award.