Prateeth Rao
MS (Research),B.Tech
Researcher in 3D Geometric Computer Vision, Remote Sensing, Applied Deep Learning and Embodied AI (VLA)
About Me
I am a Master’s student at IIIT-Bangalore with research interests spanning Machine Vision, Applied Deep Learning, 3D Vision, VLA and Geospatial Intelligence. My work at IITT-B focused on leveraging Epipolar geometry supervision to Visual Odometry (VO) by improve the robustness and reliability of modern SLAM systems, particularly under challenging real-world and poor texture environments.
I'm always excited to explore new frontiers. Currently, I'm diving into multimodal Visual SLAM (Neuromorphic/Radar based), developing hardware deployable VLA, and analysing changes in environment using Remote sensing methods.
Beyond research, I maintain a strong personal interest in Yoga, caring for plants in my garden, and music across genres including Classical, Bollywood, and EDM. I am also deeply curious about ancient knowledge systems and history. I enjoy exploring ancient architecture, studying historical texts, and understanding cultural evolution through archaeological and literary perspectives. In parallel, I take an educational interest in Ayurvedic and homeopathic traditions, particularly from the viewpoint of learning about trees, medicinal plants, and their historical significance.
Education
M.S in Data Science and AI
IIIT-B, Bengaluru (Aug 2023-Jun 2026)
Specializations: Machine Perception and 3D Reconstruction
GPA: 3.75/4.0
B.Tech in ECE
PES University, Bengaluru (Aug 2019- May 2023)
Specializations: Image Processing and Medical Informatics
GPA: 8.5/10.0
Core Skills & Interests
Languages
Python, C++, Embedded C, LaTeX
Python Tools
Numpy, OpenCV, Scipy, Tensorflow, Pytorch, Kornia, TensorRT
Other Tools
Git, ROS2, GEE, GRASS GIS, AWS
Interests
Global Scene Graphs in Loop Closures,VLA reasoning from human action videos for manipulators, Graph Based Visual Odometry, Multimodal Visual SLAM, Temporal Forecasting in GIS, Physics based DDPM for Visual Odometry and Agriculture Robotics (Combining Drone and Manipulator)
Experience
Drone Perception Research Consultant (Part time)
Haveli UAV's at IIIT-B| Sept 2023 - Dec 2024
- Deployed Monocular Depth Estimation with SLAM on ROS
- Constributed towards Semantic segmentation of tea estate
- Researched on Visual SLAM deployable methods for RPi 4
Project Intern
COMET IIIT-B, Bengaluru | Feb 2023 - Jul 2023
- Developed components of the 5G NR transmitter for the PDCCH stack.
- Researched and developed a deep learning model for CSI receiver channel parameters, achieving state-of-the-art performance.
- Published a paper under the guidance of Dr. Satish Kumar at IEEE GCITC 2023.
Senior Intern
MedInn TechLab @ PESU, Bengaluru | Sept 2022 - May 2023
- Led the capstone project "Video Analytics for User identification in Hospitals".
- Served as a team leader and lab mentor for young researchers in Medical Informatics.
- Published 6 papers in the domain of Deep Learning and Medical Imaging.
Core Team IoT
C.O.D.S, PES University | Oct 2020 - Jan 2022
- Worked on developing ADAS (Advanced Driver-Assistance Systems) through sensors and IoT cloud.
Key Projects
Building VLA model for Manipulators (Supervised Approach)
Designed and implemented and Vision-Language-Action (VLA) foundation model inte-grating a multi-modal ResNet architecture with a Large Language Model (Gemma-2B).
Formulated and implemented a custom weighted Mean Squared Error (MSE) loss function that applies a 2.0x penalty to end-effector/gripper inaccuracies, significantly improving the model’s fine-motor grasping capabilities.
View Synthesis in Unbounded Scenes using Radiance Fields
Explored the theoretical limitations of standard MLPs in capturing high-frequency spatial details, successfully implementing Positional Encoding to map spatial (x, y, z) coordinates to high-dimensional frequency domains.
Addressed the ”spectral bias” and infinite scale challenges of the KITTI and Mip-NeRF 360 datasets by applying non-linear spatial contraction formulas to normalize scene geometries.
Investigated inverse rendering techniques, deriving and implementing the discrete volumetric rendering equation (alpha-compositing and transmittance) to optimize continuous 3D density and color fields purely from 2D image supervision.
Building Visual Odometry module in CPP
Built a small orbSLAM based VO module with integration of pose graph optimization
Noise Modeling in Dense Matching with DDPM(Thesis Extension)
Modeling Noise in Dense Matching as a Isotropic Gaussian Noise problem modelled using a Simple SA based DDPM model with a weighted Graph Propagation transformer used in Relative Pose Estimate
Python V-SLAM Algorithm (Defenced Thesis)
My thesis, 'GRAPH-BASED RELATIVE CAMERA POSE ESTIMATION WITH EPIPOLAR GEOMETRY SUPERVISION FOR VISUAL SLAM', involves a deep dive into Visual SLAM. The thesis highlights learning an Approximation to 8-point method using in Pose estimation from essential matrix using graph learning methods in camera normalized space. Composite Objective loss function containing Epipolar Supervision loss, Scale Flow loss, heading angle loss and MSE of predicted quaternions and translation vectors. Minimizing objective function yields a similar minimization of Ae=0 on epipolar constraints. Compared with existing Spatial Pose regression models to highlight superiority of epipolar geometry in finding motion embeddings.
Deployment of Various Global Path Planning methods on Self-Constrained environment
Global Path planners such as Djikstra, A*, D*, RRT, RRT* are used in path planning from start node to goal node by finding the shortest path between the nodes. The project also highlights different improvements to standard path planning algorithms.
Dynamic Mobile Robot Navigation in ROS
This project involved simulating a mobile robot navigating a 2D map with both static and dynamic obstacles. I utilized ROS for control, Gazebo for physics simulation, and advanced deep learning techniques like semantic segmentation and keypoint estimation for environment perception. The research was to find 3 possible locations on the static map using semantic segmentation and later perform triangulation (find 3D depth) for navigation.
Road Width Estimation using GEE & Deep Learning
Collected and manually annotated geospatial data from Google Earth Engine (GEE) along a national highway. I trained a deep residual U-Net neural network to perform precise road segmentation from this data, enabling accurate estimation of road width—a key task in infrastructure monitoring and analysis. It was observed that while Google Earth provided a width of 6m, my method estimated 3.5m, closest to the actual ground truth of 3m.
Latest Blog Posts
Bird Chirp Signal Analysis using Handcrafted Signal Processing Features
Explained time-frequency representations, and feature extraction of bird audio signals from BirdCLEF dataset with plots and observations.
Event-based Visual Odometry and Inertial Odometry
Detailed experiments on event camera VO, robust estimation with RANSAC/MAGSAC++, and IMU fusion strategies.
Research Papers
- "EpiDiffVO: Geometry Aware Epipolar Diffusion for Robust VO" Under Revision to IEEE RA-L journal, ArXiv Preprint : arXiv:2605.19556
- "Relational Epipolar Graphs for Robust Relative Camera Pose Estimation" Submitted to IJCV: arXiv preprint arXiv: 2604.04554, Prateeth Rao and Sachit Rao.
- "Intelligent Navigation Tactics for Differential Drive Robots: Expanding Boundaries" published at IEEE CONECCT 2024.
- "Channel Estimation for mmWave MIMO using Deep Residual Learning" published at IEEE GCITC 2023.
- "Advancements in Semantic Skin Lesion Segmentation: A Comparative Study with YOLOv8 for Accurate Skin Lesion Segmentation" published at IEEE GCITC 2023.
- "Exploring FMU-NET: A Comprehensive Study on Person Identification Through Novel Semantic Segmentation Techniques" published at IEEE EASCT 2023.
- "TMV-NET: Regression approach for the quantification of Tympanic Membrane" published at IEEE CENTCON 2023. (Best Paper Award)
- "Real-Time Face classification and Shape based feature extraction with Deep Neural networks" published at IEEE 14C 2022.
- "YOLOv7 based face extraction and novel Eye feature extraction with EYENET" published at IEEE 14C 2022.
- "Vasculature Detection in Tympanic Membrane" published at IEEE 14C 2022.