Prateeth Rao

Prateeth Rao

MS (Research),B.Tech

Researcher in 3D Geometric Computer Vision, Remote Sensing, Applied Deep Learning and Embodied AI (VLA)

About Me

I am a Master’s student at IIIT-Bangalore with research interests spanning Machine Vision, Applied Deep Learning, 3D Vision, VLA and Geospatial Intelligence. My work at IITT-B focused on leveraging Epipolar geometry supervision to Visual Odometry (VO) by improve the robustness and reliability of modern SLAM systems, particularly under challenging real-world and poor texture environments.

I'm always excited to explore new frontiers. Currently, I'm diving into multimodal Visual SLAM (Neuromorphic/Radar based), developing hardware deployable VLA, and analysing changes in environment using Remote sensing methods.

Beyond research, I maintain a strong personal interest in Yoga, caring for plants in my garden, and music across genres including Classical, Bollywood, and EDM. I am also deeply curious about ancient knowledge systems and history. I enjoy exploring ancient architecture, studying historical texts, and understanding cultural evolution through archaeological and literary perspectives. In parallel, I take an educational interest in Ayurvedic and homeopathic traditions, particularly from the viewpoint of learning about trees, medicinal plants, and their historical significance.

Education

M.S in Data Science and AI

IIIT-B, Bengaluru (Aug 2023-Jun 2026)

Specializations: Machine Perception and 3D Reconstruction

GPA: 3.75/4.0


B.Tech in ECE

PES University, Bengaluru (Aug 2019- May 2023)

Specializations: Image Processing and Medical Informatics

GPA: 8.5/10.0

Core Skills & Interests

Languages

Python, C++, Embedded C, LaTeX

Python Tools

Numpy, OpenCV, Scipy, Tensorflow, Pytorch, Kornia, TensorRT

Other Tools

Git, ROS2, GEE, GRASS GIS, AWS

Interests

Global Scene Graphs in Loop Closures,VLA reasoning from human action videos for manipulators, Graph Based Visual Odometry, Multimodal Visual SLAM, Temporal Forecasting in GIS, Physics based DDPM for Visual Odometry and Agriculture Robotics (Combining Drone and Manipulator)

Experience

Drone Perception Research Consultant (Part time)

Haveli UAV's at IIIT-B| Sept 2023 - Dec 2024

  • Deployed Monocular Depth Estimation with SLAM on ROS
  • Constributed towards Semantic segmentation of tea estate
  • Researched on Visual SLAM deployable methods for RPi 4

Project Intern

COMET IIIT-B, Bengaluru | Feb 2023 - Jul 2023

  • Developed components of the 5G NR transmitter for the PDCCH stack.
  • Researched and developed a deep learning model for CSI receiver channel parameters, achieving state-of-the-art performance.
  • Published a paper under the guidance of Dr. Satish Kumar at IEEE GCITC 2023.

Senior Intern

MedInn TechLab @ PESU, Bengaluru | Sept 2022 - May 2023

  • Led the capstone project "Video Analytics for User identification in Hospitals".
  • Served as a team leader and lab mentor for young researchers in Medical Informatics.
  • Published 6 papers in the domain of Deep Learning and Medical Imaging.

Core Team IoT

C.O.D.S, PES University | Oct 2020 - Jan 2022

  • Worked on developing ADAS (Advanced Driver-Assistance Systems) through sensors and IoT cloud.

Key Projects

Building VLA model for Manipulators (Supervised Approach)

Designed and implemented and Vision-Language-Action (VLA) foundation model inte-grating a multi-modal ResNet architecture with a Large Language Model (Gemma-2B).

Formulated and implemented a custom weighted Mean Squared Error (MSE) loss function that applies a 2.0x penalty to end-effector/gripper inaccuracies, significantly improving the model’s fine-motor grasping capabilities.

View Synthesis in Unbounded Scenes using Radiance Fields

Explored the theoretical limitations of standard MLPs in capturing high-frequency spatial details, successfully implementing Positional Encoding to map spatial (x, y, z) coordinates to high-dimensional frequency domains.

Addressed the ”spectral bias” and infinite scale challenges of the KITTI and Mip-NeRF 360 datasets by applying non-linear spatial contraction formulas to normalize scene geometries.

Investigated inverse rendering techniques, deriving and implementing the discrete volumetric rendering equation (alpha-compositing and transmittance) to optimize continuous 3D density and color fields purely from 2D image supervision.

Building Visual Odometry module in CPP

Built a small orbSLAM based VO module with integration of pose graph optimization

Noise Modeling in Dense Matching with DDPM(Thesis Extension)

Modeling Noise in Dense Matching as a Isotropic Gaussian Noise problem modelled using a Simple SA based DDPM model with a weighted Graph Propagation transformer used in Relative Pose Estimate

Python V-SLAM Algorithm (Defenced Thesis)

My thesis, 'GRAPH-BASED RELATIVE CAMERA POSE ESTIMATION WITH EPIPOLAR GEOMETRY SUPERVISION FOR VISUAL SLAM', involves a deep dive into Visual SLAM. The thesis highlights learning an Approximation to 8-point method using in Pose estimation from essential matrix using graph learning methods in camera normalized space. Composite Objective loss function containing Epipolar Supervision loss, Scale Flow loss, heading angle loss and MSE of predicted quaternions and translation vectors. Minimizing objective function yields a similar minimization of Ae=0 on epipolar constraints. Compared with existing Spatial Pose regression models to highlight superiority of epipolar geometry in finding motion embeddings.

Deployment of Various Global Path Planning methods on Self-Constrained environment

Global Path planners such as Djikstra, A*, D*, RRT, RRT* are used in path planning from start node to goal node by finding the shortest path between the nodes. The project also highlights different improvements to standard path planning algorithms.

Dynamic Mobile Robot Navigation in ROS

This project involved simulating a mobile robot navigating a 2D map with both static and dynamic obstacles. I utilized ROS for control, Gazebo for physics simulation, and advanced deep learning techniques like semantic segmentation and keypoint estimation for environment perception. The research was to find 3 possible locations on the static map using semantic segmentation and later perform triangulation (find 3D depth) for navigation.

Road Width Estimation using GEE & Deep Learning

Collected and manually annotated geospatial data from Google Earth Engine (GEE) along a national highway. I trained a deep residual U-Net neural network to perform precise road segmentation from this data, enabling accurate estimation of road width—a key task in infrastructure monitoring and analysis. It was observed that while Google Earth provided a width of 6m, my method estimated 3.5m, closest to the actual ground truth of 3m.

Latest Blog Posts

Research Papers