mirror of
https://github.com/introlab/rtabmap.git
synced 2026-10-05 17:47:49 +08:00
updated examples in readme
This commit is contained in:
@@ -27,7 +27,7 @@ For every node:
|
||||
2. **Lidar edges.** The scan is projected in the camera at a coarse resolution (`--decimation`, default 1/4 of the image), where the projected points are dense, keeping the nearest point per cell so that surfaces hidden from the camera do not count. A cell's point is a lidar edge point when:
|
||||
- it is in front of a depth discontinuity: a neighboring cell is at least 15% farther (`--jump`) than where this cell's surface would continue, as at the border of an object in front of a farther background. The continuation is predicted from the opposite neighbor (on a plane, inverse depth is linear in the image), so that a surface seen at a grazing angle, such as the floor ahead of a low camera, is not taken for a discontinuity: its depth changes fast from one cell to the next, but as its continuation predicts; or
|
||||
- it is on an intensity edge: the lidar's intensity changes by more than 40% to a neighboring cell (`--intensity_jump`), as at a change of paint or material. A cell's intensity is the mean (of the logarithm) over all the points of the surface it sees, not its nearest point's: a lidar's beams do not return the same intensity from the same surface, and a node's scan assembles several sweeps, so cells seen by different beams would otherwise differ and make false edges along the beams' traces, even on a flat uniform wall. Then it is median filtered against what speckle remains. As for creases (below), the cells only say where there is an intensity edge: where exactly is found at the image's full resolution. This is what finds edges on flat surfaces, where the depth does not jump: holds on a climbing wall, window frames, panels. `--no_intensity` turns it off; or
|
||||
- it is on a crease: the surface normals of neighboring cells differ by more than 45 degrees (`--crease_angle`, 0 disables), as between a wall and the floor, where the depth does not jump. Normals are computed on the scan voxelized (`--crease_voxel`, 10 cm), smoother than at full resolution; each voxel's normal is spread over the cells it covers, and averaged per cell like intensity. As the normals then change over a band of a few cells across a crease (and intensity across an intensity edge), where exactly it is is found at the image's full resolution, as Canny finds edges: the normals (or intensity) are interpolated and smoothed, and the edge is where their change is the largest across it, to a fraction of a pixel. Its depth is interpolated there (these edges are on continuous surfaces), and the point is put back in 3D, one per cell. Taking the nearest lidar point of each cell instead would put them anywhere in a band a few cells wide on both sides of the edge. On the example below, creases were about as reliable as intensity edges, but did not change the result: the intensity edges already gave the same constraints there. They should help more where surfaces all reflect alike.
|
||||
- it is on a crease: the surface normals of neighboring cells differ by more than 45 degrees (`--crease_angle`, 0 disables), as between a wall and the floor, where the depth does not jump. Normals are computed on the scan voxelized (`--crease_voxel`, 10 cm), smoother than at full resolution; each voxel's normal is spread over the cells it covers, and averaged per cell like intensity. As the normals then change over a band of a few cells across a crease (and intensity across an intensity edge), where exactly it is is found at the image's full resolution, as Canny finds edges: the normals (or intensity) are interpolated and smoothed, and the edge is where their change is the largest across it, to a fraction of a pixel. Its depth is interpolated there (these edges are on continuous surfaces), and the point is put back in 3D, one per cell. Taking the nearest lidar point of each cell instead would put them anywhere in a band a few cells wide on both sides of the edge. They help most where surfaces all reflect alike, with few intensity edges.
|
||||
|
||||
Each point is weighted by the size of its jump. The edge points are kept in 3D, in the scan's frame, so that they can then be projected at full resolution.
|
||||
|
||||
@@ -40,7 +40,7 @@ The least-squares solvers (if RTAB-Map is built with g2o or GTSAM) minimize inst
|
||||
|
||||
- `g2o`, `gtsam`: Levenberg-Marquardt, with a Welsch robust kernel, which makes a point far from any image edge count for nothing, as the score does: many lidar edges have no counterpart in the image. Levenberg-Marquardt only follows the local slope, and each point is pulled toward its nearest image edge, often not its own when far from the solution, so on its own it stops in a local minimum a few degrees away. A coarse `pattern` search (2 down to 0.25 degree) gets close first, then the kernel's scale goes from wide to narrow (9, 3, then 1 x `--sigma`). Several times longer than `simplex`, and about 2 GB of memory for 500k lidar edge points.
|
||||
|
||||
The lidar edge points are then selected again with the result, and the solver runs a second time from there. On the example below, all four solvers agree within 0.07 degree.
|
||||
The lidar edge points are then selected again with the result, and the solver runs a second time from there.
|
||||
|
||||
### Code
|
||||
|
||||
@@ -65,22 +65,34 @@ Only the rotation is estimated unless `--translation` is given. The translation
|
||||
|
||||
### Checking it
|
||||
|
||||
The tool prints, for each kind of lidar edge, the share of its points within 2 pixels of an image edge once corrected: how much each brings, and how much of it is noise (on the example below, 42% of the depth discontinuities and 38% of the intensity edges). Then:
|
||||
The tool prints, for each kind of lidar edge, the share of its points within 2 pixels of an image edge once corrected: how much each brings, and how much of it is noise. Then:
|
||||
|
||||
- **Each half of the nodes on its own.** The even and the odd nodes are calibrated separately. If they agree, the result is supported by the data; if they differ, the difference is about how much the result can be trusted.
|
||||
- **The sensitivity**, with `--verbose`: how much the score drops, in percent, with the result off by 1 degree (roll, pitch, yaw) or 2 cm (x, y, z) along or about each axis of the camera, the mean of both directions. Think of the score as a valley with the result at its bottom: the sensitivity is how steep its sides are along each axis.
|
||||
- A large drop (several percent or more for 1 degree) means the data determines that axis well: a small error on it would misalign many edges, so the solver cannot be far off, and a correction on that axis can be trusted, however large.
|
||||
- Almost none (around 1% or less) means the axis is not observable from this data: the edges hardly move with it, so any value the solver finds on it, large or small, is not reliable.
|
||||
- It says how sure the result is, not how far off the camera was: that is the correction itself. It depends on the scene and the sensors, not on the error: e.g., many vertical and horizontal edges make yaw and pitch steep, roll (about the viewing axis) moves edges little near the image center so it is usually less steep, and translation is flat unless surfaces are close (a 2 cm shift moves edges a fraction of a pixel at several meters).
|
||||
- On the netherdrone data of the example below, once its mast angle was fixed: roll 7%, pitch 16%, yaw 20%, while x, y and z are 0.5 to 2%: the rotation is well determined, the translation is not (hence rotation only by default).
|
||||
|
||||
With `--images dir`, the image of every node is saved darkened, with the image edges the alignment uses (Canny) in green and the lidar edge points projected over them, depth discontinuities in red, creases in blue and intensity edges in yellow, before (`<id>_edges_1_before.png`, from where the search started: the database's camera transform, with `--initial_rotation` if given) and after (`<id>_edges_2_after.png`) the correction, with the node's id and speeds in the top left corner: the mean since the previous node (over which an assembled scan is taken) and the instantaneous one when the node was added. After, the lidar points should lie on green wherever both sensors see an edge. Points away from any green, and green edges without points, are edges only one of the sensors sees (e.g., lidar intensity through glass, or shadows in the image).
|
||||
|
||||
For example, a node of an indoor climbing gym, before (left) and after (right) the correction: before, the creases along the floor and the climbing holds' outlines are off their green edges; after, they lie on them.
|
||||
|
||||
| Before (`<id>_edges_1_before.png`) | After (`<id>_edges_2_after.png`) |
|
||||
|---|---|
|
||||
|  |  |
|
||||
|
||||
With each of them, two images show the maps the lidar edges are found from, at the decimated resolution, half transparent over the image, with the image's edges in green and the map's own edges (Canny, as for the image) in blue:
|
||||
|
||||
- `<id>_intensity_1_before.png`, `<id>_intensity_2_after.png`: the lidar's intensity per cell, as used (mean log-intensity, median filtered), from red (dark) to yellow (bright), with the contrast stretched for each image. The intensity edges are where it changes; they should line up with green where the image shows the same change of material.
|
||||
- `<id>_normals_1_before.png`, `<id>_normals_2_after.png`: how each cell's surface faces the camera, from yellow (facing it) to red (seen edge on, at a grazing angle). Surfaces at a grazing angle are where depth changes fast without a discontinuity, and where the intensity drops. The normals are computed on the scans voxelized at `--crease_voxel`, so this image is made whether or not creases are used.
|
||||
|
||||
For the same node, before (left) and after (right) the correction: the holds stand out in yellow in the intensity, as do the walls' and floor's orientations in the normals; once corrected, the map's edges (blue) lie on the image's (green).
|
||||
|
||||
| Before | After |
|
||||
|---|---|
|
||||
|  |  |
|
||||
|  |  |
|
||||
|
||||
It is also worth running it again with other values of `--decimation`, `--intensity_jump` or `--voxel` (which voxel filters the scans first, to compare densities): a result that does not move with them is more trustworthy than one that does.
|
||||
|
||||
## Options
|
||||
@@ -94,7 +106,7 @@ It is also worth running it again with other values of `--decimation`, `--intens
|
||||
| `--decimation #` | `4` | Image decimation at which the lidar edges are found. Lower is finer, but needs denser scans. |
|
||||
| `--jump #.#` | `0.15` | Relative depth jump for a depth discontinuity, as a fraction of the point's depth: 0.15 means a neighbor at least 15% of its depth farther than where its surface would continue (e.g., 0.6 m behind a point at 4 m). |
|
||||
| `--intensity_jump #.#` | `0.4` | Relative intensity change for an intensity edge, as a fraction: 0.4 means a neighbor at least 1.4 times brighter or darker (compared on log-intensity). Lower finds more edges, and more speckle. |
|
||||
| `--intensity_weight #.#` | `1` | Weight of the intensity edges relative to the depth discontinuities. Lower it where intensity is less reliable than geometry (e.g., much glass); on the example below, lowering it made the two halves of the nodes agree less, most on the tilt (likely because the intensity edges add horizontal edges, which the depth discontinuities there have fewer of). |
|
||||
| `--intensity_weight #.#` | `1` | Weight of the intensity edges relative to the depth discontinuities. Lower it where intensity is less reliable than geometry (e.g., much glass). |
|
||||
| `--crease_angle #.#` | `45` | Creases: where the surface normals differ by this angle (deg). 0: off. |
|
||||
| `--crease_voxel #.#` | `0.1` | Voxel size (m) of the scans on which the normals are computed. |
|
||||
| `--no_intensity` | | Use only depth discontinuities. |
|
||||
@@ -113,25 +125,6 @@ It is also worth running it again with other values of `--decimation`, `--intens
|
||||
- **Intrinsics.** The camera's calibration (focal lengths, center, distortion) is assumed right; an error there biases the rotation.
|
||||
- **One camera per node.** Nodes with several cameras are skipped.
|
||||
|
||||
## Example
|
||||
|
||||
On the netherdrone demo of rtabmap_ros (an Ouster OS1-32 on a rotating mast and a camera, 80 nodes, full resolution scans), with the default `simplex` solver:
|
||||
|
||||
```
|
||||
Consistency, each half of the nodes on its own:
|
||||
even nodes: xyz=(0.000, 0.000, 0.000) m rpy=(2.78, -0.72, -0.46) deg
|
||||
odd nodes: xyz=(0.000, 0.000, 0.000) m rpy=(3.07, -0.38, -0.55) deg
|
||||
|
||||
Correction of the camera's mount X, in the camera's body frame (x forward, y left, z up):
|
||||
xyz (m): 0.000000 0.000000 0.000000
|
||||
roll pitch yaw (rad): 0.052874 -0.009342 -0.008714 (deg: 3.0294 -0.5353 -0.4993)
|
||||
quaternion (x y z w): 0.026413 -0.004785 -0.004232 0.999631
|
||||
```
|
||||
|
||||
That is a roll of about 3 degrees of the camera about its viewing axis, and about half a degree of tilt and of pan, the same for both halves within about 0.3 degree. The translation was not observable (a flat score), and was kept as measured.
|
||||
|
||||
Most of that roll turned out not to be the camera's: the lidar's mast angle was published with a scale error (one count per motor turn), which turned the lidar about the mast axis, the camera's viewing axis, more and more over the session. Calibrating each node on its own showed it: their roll drifted by about 4 degrees from the first node to the last. With the mast angle fixed, the camera's roll was 1.35 degrees.
|
||||
|
||||
## Reference
|
||||
|
||||
J. Levinson and S. Thrun, "Automatic Online Calibration of Cameras and Lasers", *Robotics: Science and Systems IX* (RSS), 2013. [PDF](http://www.roboticsproceedings.org/rss09/p29.pdf)
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 143 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 136 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 153 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 142 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 140 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 136 KiB |
Reference in New Issue
Block a user