AI-Enhanced Autonomous Mobile Robot for Industrial Applications

previewfull image
Project Summary: Developed an autonomous mobile robot integrating YOLOv8 object detection with ROS 2 navigation. Using a TurtleBot3 Burger and Raspberry Pi 4B, the project explored embedded AI for industrial robotics through simulation, model evaluation, and real-world deployment.

Introduction

How capable can a low-cost mobile robot be when asked to run modern deep-learning models? That was the question behind my capstone project: developing an autonomous mobile robot for industrial applications using AI-based object detection and navigation.

I independently developed and evaluated the system using a TurtleBot3 Burger powered by a Raspberry Pi 4B, integrating YOLOv8 object detection with ROS 2 and Gazebo simulation. The robot was designed to navigate a maze, identify four target classes apples, bananas, oranges, and a person and approach them in a prescribed order.

From creating a custom object-detection dataset to evaluating hardware platforms, benchmarking models, integrating navigation, and testing the physical robot, I investigated the practical challenges of deploying AI on resource-constrained robotic hardware.

The project ultimately explored a broader engineering question: how do hardware limitations, software compatibility, model performance, and real-world conditions influence the feasibility of an autonomous robotic system?

The Problem

Modern object detection models can be remarkably capable, but that capability usually comes with a computational cost. For a desktop or cloud-based application, that cost may be acceptable. For a small mobile robot running on a Raspberry Pi, it becomes a significant constraint. The goal was therefore not simply to achieve the highest possible object-detection accuracy.

Instead, the project asked:

Can YOLOv8 provide useful object detection on a low-cost embedded robotics platform, and can that detection be integrated into a navigation task?

This required balancing several competing requirements:

To investigate these trade-offs, I divided the project into two stages: simulation and physical testing.

The Robot

The physical platform used for the project was the TurtleBot3 Burger, a small and relatively inexpensive mobile robot designed around the Raspberry Pi and ROS ecosystem.

The robot provided the basic mobile platform, while cameras were used to provide visual information to the object-detection system.

previewfull image
One of the first challenges was determining which Raspberry Pi could actually support the required workload.

I tested three generations of Raspberry Pi:

The expectation was fairly straightforward: newer hardware should provide better performance.

The reality was more complicated.

The Raspberry Pi 3B+ was unable to provide a usable platform for YOLOv8. The Raspberry Pi 5 offered considerably more computational capability, but firmware compatibility issues prevented it from being a practical solution for the complete ROS 2 system.

The Raspberry Pi 4B ultimately provided the most stable platform for the project.

This was an important early lesson: the fastest hardware on paper is not necessarily the best hardware for a complete system.

Compatibility, stability and software support can matter just as much as raw computational performance.

Building the Detection System

For object detection, I selected YOLOv8. YOLO, or You Only Look Once, is designed to perform object detection in a single inference process, making it particularly attractive for applications where detection speed is important. YOLOv8 is available in several model sizes, ranging from the lightweight nano model to substantially larger models.

I evaluated:

The larger models generally offer greater capacity, but that comes at the cost of increased computational requirements.

On a desktop GPU, this trade-off can be relatively easy to manage. On a Raspberry Pi mounted to a small mobile robot, it becomes much more restrictive. The models were deployed as ROS 2 nodes so that the object-detection system could communicate with the rest of the robotic system.

This allowed the project to move beyond simply asking whether YOLOv8 could detect an object and instead ask whether its output could actually contribute to the robot’s behaviour.

Creating the Dataset

In addition to evaluating pretrained models, I created a custom dataset containing 570 images of the target objects.

The dataset was used to investigate whether tailoring the model to the specific environment and objects would improve its performance.

The target classes were:

The models were evaluated using several common object-detection metrics, including:

This provided a way of comparing not just whether the model could detect an object, but how accurately and efficiently it could do so.

previewfull image

The Hardware Constraint

One of the biggest findings was that the hardware quickly became the limiting factor. The larger YOLOv8 models were simply not practical on the available Raspberry Pi hardware.The nano and small variants were the only models that provided a realistic path towards deployment. Even then, the system was far from what would normally be considered real-time computer vision. The feasible configurations were operating at approximately one frame every six seconds. That number sounds extremely slow compared with the frame rates normally associated with computer vision.

However, it also demonstrates one of the central challenges of embedded AI:

A model doesn’t have to be fast in isolation. It has to be fast enough for the task it is being used for.

For this particular project, the robot did not need to identify hundreds of objects per second. It needed to identify objects well enough to make navigation decisions within the constraints of the environment. That distinction made the smaller models much more useful than their larger counterparts.

previewfull image

Accuracy vs Performance

The results demonstrated the expected trade-off between model size, accuracy and computational requirements. Of the models tested, the nano and small variants were the most viable for deployment. The pretrained YOLOv8-small model produced the best overall balance between accuracy and speed, achieving an mAP of 81%.

This made it the strongest candidate for the physical robot. Interestingly, the custom-trained models produced strong precision but suffered from reduced recall. Even more unexpectedly, some of the custom-trained models were slower despite having smaller file sizes. This was an important result because it challenged a simple assumption:

A smaller model file does not necessarily mean faster inference.

Inference performance depends on considerably more than the size of the model on disk. The underlying architecture, operations performed during inference, hardware optimisation and software implementation all contribute to the actual execution time. This was one of the more useful lessons from the project because it reinforced the importance of measuring performance on the target hardware rather than relying on theoretical specifications.

From Simulation to Reality

Before testing the complete system on the physical robot, I created a simulated maze using Gazebo. This provided a controlled environment in which the detection and navigation system could be tested without the unpredictability of the physical robot.Simulation offered several advantages. It was easier to repeat experiments, modify the environment and test different configurations without worrying about damaging hardware or resetting the robot. However, the simulation also introduced an important question:

Would the system that worked in simulation actually work on the physical robot?

The answer was: mostly, but not in exactly the same way.

previewfull image

The Sim2Real Gap

The physical experiments were conducted in a laboratory maze designed to replicate the simulated environment. The results showed a significant difference between the two environments. In simulation, navigation trials typically took approximately 150–170 seconds. In the physical environment, successful trials were considerably faster, generally taking around 70–100 seconds. At first, this was surprising. The physical robot was completing the task faster than the simulated robot, but the physical results were also less consistent.

This highlighted an important characteristic of robotics development: simulation does not perfectly reproduce reality. The simulated environment provides controlled conditions, while the physical robot introduces factors such as:

These differences contribute to what is commonly described as the Sim2Real gap. The project demonstrated that a system can behave predictably in simulation while still producing significantly different results when deployed in the real world. Simulation was therefore extremely useful, but it could not replace physical testing.

previewfull image

Integrating Detection and Navigation

The most important part of the project was integrating all of these components into a single system.

The basic process was:

Camera → YOLOv8 → Object Detection → ROS 2 → Navigation → Robot Movement

The camera provided visual information. YOLOv8 processed that information and identified objects. The detection results were passed through the ROS 2 system, where they could be used alongside the robot’s navigation system.

The robot then used the available information to navigate through the maze and locate the required targets in their prescribed order. This integration was significantly more complicated than simply running an object-detection model.

A model can achieve an impressive accuracy score and still be unsuitable for a robot if its inference time is too slow, if detections are inconsistent, or if the output cannot be reliably incorporated into the navigation system. The project therefore became an exercise in systems engineering as much as AI.

previewfull image

What Didn’t Work

Some of the most useful results came from the approaches that failed. The Raspberry Pi 3B+ was unable to provide a suitable platform for YOLOv8. The Raspberry Pi 5, despite offering greater computational capability, was affected by firmware compatibility issues that prevented it from becoming the preferred platform. The larger YOLOv8 models exceeded the practical capabilities of the embedded hardware.

Custom training also did not automatically produce better results. While some custom models achieved strong precision, recall decreased and inference performance was unexpectedly slower. These failures were not simply problems to be removed from the final result. They demonstrated why embedded AI deployment requires testing the complete system rather than optimising individual components independently.

previewfull image

What I Learned

The biggest lesson from this project was that deploying AI on an embedded robot is fundamentally an integration problem.

It is tempting to think of the problem as:

Choose the best AI model → put it on the robot → make the robot use it.

In practice, the process is much more complicated. The best model on a powerful computer may not be the best model for the robot. The newest Raspberry Pi may not be the best hardware if the required software stack is unstable.

A smaller model may not necessarily be faster. A model with higher precision may not be more useful if its recall is too low. And a system that works perfectly in simulation may behave differently in the physical environment. The project therefore reinforced the importance of evaluating AI systems according to their actual deployment requirements. Accuracy matters, but so do inference time, hardware compatibility, stability, reliability and integration.

Where I Would Take It Next

There are several areas where the system could be improved. The most obvious would be improving inference speed. Running detection at approximately one frame every six seconds places a significant limitation on how quickly the robot can react to changes in its environment. Hardware acceleration or a dedicated AI accelerator could potentially improve this substantially. The dataset could also be expanded to include greater variation in lighting, object orientation, distance and background conditions. This could improve robustness and potentially reduce the gap between controlled training data and real-world environments.

Navigation could also be improved by more tightly integrating object detection with localisation and path planning rather than treating detection primarily as a separate perception stage. Finally, more extensive physical testing would be required before considering the system suitable for a genuine industrial environment. The project demonstrated feasibility, but a laboratory prototype is very different from a production-ready autonomous robot.

Conclusion

This project began with a relatively simple question:

Can a low-cost Raspberry Pi-powered robot run modern deep-learning object detection and use it to navigate a real environment?

The answer is yes, but with significant limitations.

The Raspberry Pi 4B provided the most stable hardware platform, while YOLOv8-small offered the best balance between accuracy and computational requirements. The pretrained model achieved an mAP of 81%, while the larger YOLOv8 variants were not practical for the available hardware.

The physical experiments also demonstrated that successful simulation does not guarantee identical real-world performance. The difference between the simulated and physical navigation results highlighted the continuing challenge of transferring robotic systems from controlled environments into reality. More importantly, the project demonstrated that embedded AI is not simply a matter of selecting the most accurate model.

The real challenge is finding a combination of hardware, software, model architecture and algorithms that works together within the constraints of the physical system.

That is ultimately what made this project valuable to me. Rather than simply building an AI model or a mobile robot, I had to bring the two together and discover where the boundaries of the system actually were.