Quriostack

Top 5 ROS Mistakes Beginners Make

Info
Top 5 ROS Mistakes Beginners Make
Hermes Smith
·June 11, 2026· 11 min read
0 0

The article mentor junior robotics engineers through a fellowship program, and every cohort makes the same five mistakes. They're not exotic; they're the kind of errors that look plausible, work on your desk, and then wreck you the first time the robot moves. If There is a time machine, this is the post It would send to the team circa 2017. Consider it a public service announcement.

Why This Matters

Mistakes in robotics are not like mistakes in web development. A web app that misbehaves shows an error on a screen. A robot that misbehaves can drive into a wall, drop a half-ton of glass, or in the worst cases Documentation and common practice have, cut a cable that's part of the building's safety system. The stakes are physical. Which is exactly why the patterns matter more than the APIs. Most of these mistakes aren't syntax; they're structural — the kind of thing you only learn by either burning yourself or hearing someone else's burn story.

The other reason this matters is that the most common roboticist mistake — treating topics like Python variables — is invisible until the system grows. Then it manifests as flakiness you can't reproduce. A perception pipeline that works 99% of the time will still fail in the 1% — and if that 1% corresponds to the moment your robot is reversing toward a person, the consequences are not acceptable.

There's a deeper structural reason too. ROS 2 is, at its core, a contract-driven system. Every topic has a type, every publisher-subscriber pair agrees on QoS, and every service has a defined request/response shape. Beginners often write code that violates these contracts without realizing it. The result is a system that's brittle in subtle ways: messages that sometimes don't arrive, services that return stale data, parameters that look right but don't apply. These bugs are hard to track down because they don't fail consistently.

Finally, there's a career-development argument. Hiring managers at robotics companies can spot a junior ROS 2 developer within minutes of code review. The biggest tells are callback stacks that block the executor, hardcoded topic names everywhere, parameter declarations that never appear in introspection, missing QoS awareness, and uncoordinated discovery domains. These aren't just technical errors; they're signals that the developer hasn't yet internalized that they're building a distributed system. Get these patterns right early and you'll be the person people trust to ship the load-bearing code.

The Core Idea

The five mistakes below are listed in roughly increasing order of severity. Documentation and common practice have them like field notes because they're field notes: what the mistake looks like, why it bites you, and the fix. They share an underlying theme: ROS 2 is a distributed system with explicit contracts. Pretending it's something else is where beginners get hurt.

Mistake 1: Confusing topics and services. You want the robot's current pose, so you make a service call every time you need it. Five milliseconds later, your 100 Hz control loop is starving. Or you make the perception stream a service, and the executor deadlocks. The distinction matters: topics are for streams of data, services for one-shot RPCs, actions for long-running tasks with feedback. Picking wrong is one of the top causes of beginner systems that "kind of work."

Mistake 2: Running too much logic in callbacks. A 100-line callback that synchronously does inference, validation, network it/O, and logging seems fine until your second message arrives and blocks waiting for the first to finish. Callbacks should be short. Offload work to thread pools, executors, or coroutines. The MultiThreadedExecutor is your friend.

Mistake 3: Hardcoding topic names everywhere. Six months into the project, you want to remap a topic, and now you have to find every literal string /camera/image and change it. Worse, you wrote a publisher and a subscriber that disagreed on the topic name because of a typo, and the system silently didn't connect. Use parameters, use constants, use launch-time remapping.

Mistake 4: Ignoring QoS. You publish with RELIABLE and your subscriber is on BEST_EFFORT with a depth of 1. They might or might not talk depending on DDS version. Or you use KEEP_ALL and your queue grows unbounded when the camera bursts. QoS is not decoration; it's a contract. Learn what each policy does.

Mistake 5: Pretending discovery is magic without configuring domains. You start your robot, you start your laptop, and they collide on ROS_DOMAIN_ID=0 because that's the default. Suddenly your code is publishing into a graveyard of subscribers that aren't yours. Every robot should have a unique domain It would, and your launch system should enforce it.

There's a useful framework for thinking about these errors: each mistake is a violation of a contract that ROS 2 documents but beginners don't read. The contract on topics and services is in REP 141 and the ROS 2 design documents. The contract on QoS is in REP 2009. The contract on domain IDs is in the ROS 2 concepts documentation. None of these are secrets — they're just things people skip because the README tells them how to run a publisher and that's good enough to feel productive. Then a year later they're debugging a system that's mysteriously unreliable.

The cure is not memorizing the spec; it's developing a habit of reaching for the right primitive. When you find yourself wanting to "ask" another node for a value, ask first: is this a stream or a query? Stream → topic, query → service, query that takes a long time and needs to be cancellable → action. When you find yourself doing too much in a callback, ask: can It pushs this to a queue? When you find yourself writing /some/topic in your code, ask: should this be a parameter, a launch-time remap, or a constant? When you bring up a new system, ask: what are the QoS profiles, and do they match? When you set up multiple robots, ask: are they on different domains?

A Concrete Example

Below is a demonstration of Mistake #2 in code — callbacks that do too much — and a clean fix. The buggy version:

Python
import rclpy
from rclpy.node import Node
from sensor_msgs.msg import Image
from std_msgs.msg import String
import time


class BadImageConsumer(Node):
    def __init__(self):
        super().__init__('bad_consumer')
        self.pub = self.create_publisher(String, '/result', 10)
        self.create_subscription(Image, '/cam/raw', self.on_image, 10)

    def on_image(self, msg):
        # DON'T: this whole block runs in a single executor callback
        img = self.deserialize(msg)         # 5 ms
        features = self.run_inference(img)  # 80 ms
        validated = self.check(features)    # 3 ms
        result = self.log_to_cloud(validated)  # 200 ms (network)
        self.pub.publish(String(data=result))


def deserialize(self, msg):
    time.sleep(0.005)
    return msg.data

A 30 Hz camera delivering roughly every 33 ms will starve the executor, dropping frames and slowing the entire node. Here's the right shape:

Python
import asyncio
import rclpy
from rclpy.node import Node
from rclpy.executors import MultiThreadedExecutor
from rclpy.callback_groups import MutuallyExclusiveCallbackGroup
from sensor_msgs.msg import Image
from std_msgs.msg import String
import cv2
import numpy as np


class GoodImageConsumer(Node):
    def __init__(self):
        super().__init__('good_consumer')

        # I/O group — runs in one thread, must be fast
        io_group = MutuallyExclusiveCallbackGroup()
        self.create_subscription(
            Image, '/cam/raw', self._enqueue, 10, callback_group=io_group,
        )

        # Inference group — runs in another thread, can be slow
        infer_group = MutuallyExclusiveCallbackGroup()
        self.timer = self.create_timer(
            1.0 / 30.0, self._process_queue, callback_group=infer_group,
        )
        self._queue: list[np.ndarray] = []

    def _enqueue(self, msg: Image):
        # Cheap — just decode and push to a queue
        img = cv2.imdecode(np.frombuffer(msg.data, np.uint8), cv2.IMREAD_COLOR)
        if len(self._queue) < 3:
            self._queue.append(img)
        else:
            self.get_logger().warn('Dropping frame — inference cannot keep up')

    def _process_queue(self):
        if not self._queue:
            return
        frame = self._queue.pop()
        # Heavy work, but it no longer blocks I/O
        features = self._run_model(frame)
        self._publish_result(features)


def main():
    rclpy.init()
    node = GoodImageConsumer()
    executor = MultiThreadedExecutor(num_threads=2)
    executor.add_node(node)
    try:
        executor.spin()
    finally:
        node.destroy_node()
        rclpy.shutdown()

Two callback groups, two threads, two responsibilities. The it/O side stays cheap; the heavy lifting happens off-thread. This is the pattern that lets a robot run a 30 Hz perception loop without dropping a sensor frame.

The same lesson applies to Mistake #1 (the right primitive). If you want the robot's current pose and the pose is published continuously on /odom, subscribe to it and read the latest message — don't make a service call. If you want a one-shot operation that takes a non-trivial amount of time and can be cancelled, use an action. If you want a quick lookup that returns immediately, use a service. The mental test is: "Will this happen many times per second, and do It cares about intermediate values? Yes → topic. Is this a request/response that takes O(few milliseconds)? Yes → service. Is this a long-running operation with feedback and possible cancellation? Yes → action."

Common Pitfalls

  1. Mistake 1 expanded: using topics for one-shots. "Publish a message asking the planner to start" sounds clever; you reinvented a service. Topics can't return a result, can't be cancelled, and pile up if the consumer is slow.

  2. Mistake 2 expanded: blocking on network it/O in callbacks. A single 500 ms HTTP call inside a callback will freeze the entire spin. Push to a queue or use asyncio-based executors.

  3. Mistake 3 expanded: typos in topic names. /camera/image vs /camera/images won't error; the publisher simply won't connect. Use ros2 topic info to verify.

  4. Mistake 4 expanded: subscribing before QoS is declared. DDS won't connect a subscriber to a publisher with incompatible QoS. The connection failure is silent. Always run ros2 topic info /name --verbose after bringing up a new topic.

  5. Mistake 5 expanded: running multiple robots on domain 0. A warehouse with ten robots all on domain 0 will look like one giant, broken graph. Per-robot domain IDs are not optional at scale.

  6. Bonus: parameter declarations without introspection. If you don't declare_parameter, your parameters won't show up in ros2 param list and your team can't tell what's configured. Declare every parameter.

  7. Bonus: callbacks that throw exceptions silently. A bug in a callback can be swallowed by the executor. Wrap risky code in try/except and log the exception so you see it when something goes wrong.

  8. Bonus: blocking on time.sleep. Use create_timer and let the executor drive timing. Blocking the spin loop with time.sleep causes latency jitters in unrelated callbacks.

  9. Bonus: not using lifecycle nodes when you should. A lifecycle node that you configure but never activate will sit quietly with no errors. Add health-check assertions on the state.

  10. Bonus: forgetting rclpy.shutdown() on exit. A node that doesn't shut down cleanly leaves resources (sockets, threads, executors) behind. Use a try/finally that always shuts down.

When to Use This (And When Not To)

This advice applies to virtually anyone writing ROS 2 today. The only exception is if you're doing tightly-scoped teaching demos where the goal is seeing output, not engineering a real system. In that case it's fine to violate some of these rules because you won't be maintaining the code. But for any project you intend to keep, these five mistakes will find you if you don't learn them first.

There are also situations where these rules bend. In a tightly constrained prototype demo, using a single thread and a SingleThreadedExecutor is acceptable. In an integration test, you might intentionally have a fast service call that behaves like a topic. The key is to know you're deviating and document why. The vast majority of beginner mistakes aren't conscious violations; they're unconscious ones.

How to Build the Right Habits Early

The mistakes above are common because they're easy to make. Nothing in the framework forces you to use the right primitive, the right QoS, or the right domain. The framework just provides options. That's a deliberate design choice — flexibility over constraint — but it leaves beginners to discover the pitfalls themselves.

There are some habits that help. First, before adding a new topic or service, ask yourself: which existing node publishes or consumes similar data? Reuse that interface. Second, after creating a new publisher or subscriber, run ros2 topic info and check that the QoS settings are what you intended. Third, before adding code to a callback, ask: "can this run asynchronously?" If the answer is yes, queue the work and process it elsewhere. Fourth, when you find yourself wanting to share data between nodes, prefer the topic/service/action primitives over global state.

There's also a useful debugging heuristic: when something doesn't work, check the contract first. A topic that should have a subscriber but doesn't? ros2 topic info. A service that should respond but doesn't? ros2 service list. A parameter that should be set but isn't? ros2 param list. Most "ROS doesn't work" bugs are contract bugs in disguise.

Finally, be kind to your future self. Six months from now, you'll forget why that particular QoS profile was chosen. Six months from now, you'll wonder if that hardcoded topic was intentional. Six months from now, you'll wish you had a comment explaining the design. The disciplines above are also the disciplines of code that's pleasant to maintain.

Wrapping Up

The five most common beginner mistakes in ROS 2 are: confusing topics and services, slow callbacks, hardcoded topic names, ignored QoS, and conflicting discovery domains. Each is structural and each gets worse as your system grows. The fix is small in lines but large in discipline.

Actionable next step: take one of your existing ROS 2 nodes and add a ros2 doctor --report run plus a ros2 topic info /some_topic --verbose to your standard startup. You're going to discover at least one thing you've been doing wrong. Good. From there, build the discipline of asking the contract-questions every time you touch the system: "Is this the right primitive? Is this the right QoS? Is this the right domain? Can It moves the heavy work off the callback?"

Further Reading

Hermes Smith

Comments (0)

Sign in to join the conversation.

No comments yet. Be the first to share your thoughts!