A Different Approach to Robotic Bin Picking
Industrial bin picking has traditionally depended on fixed 3D cameras, predefined grasp points, detailed calibration, and relatively controlled bin positions. These methods can perform well in structured applications, but their performance can become less predictable when parts are randomly oriented, partially occluded, or when the bin itself changes position.
Inbolt takes a different approach by placing a 3D vision camera directly on the robot arm. Instead of relying on a fixed view of the workspace, the robot can move its vision system with the end effector and continuously update its understanding of the workpiece.
In my view, this architecture is particularly significant because the camera becomes part of the robot's motion system rather than remaining a separate sensing device. That changes how the robot can respond to variations in the physical workspace.
On-Arm Vision for Unstructured Environments
The Inbolt system uses proprietary AI models together with real-time 3D vision to identify suitable grasp opportunities. The system does not depend on a single predefined grasp point. It can evaluate different accessible surfaces and determine how the robot can approach the component.
This is useful when components overlap, rotate, or expose different surfaces from one cycle to another. Rather than requiring every possible orientation to be explicitly programmed, the vision system evaluates the current scene and generates an appropriate grasp strategy.
The reported production performance is less than 1 second of processing time per pick, with pick success rates of up to 95% in live manufacturing environments.
Three-Step Closed-Loop Operation
The solution follows a continuous perception and correction process rather than treating picking as a sequence of isolated programmed movements.
1. Identify a Pickable Surface
The 3D vision system analyzes the randomly arranged components and identifies an accessible area for gripping. The robot can select different grasp approaches according to the visible geometry of the part.
2. Localize the Part After Gripping
After the robot grasps the component, the vision system can analyze the object again and determine its actual position and orientation. This provides additional information after the initial pick.
3. Correct the Placement Motion
While the robot moves toward the destination, the AI can continuously refine the trajectory based on the updated object position. This in-hand localization capability helps compensate for variations introduced during gripping.
This closed-loop architecture is, in my opinion, one of the more important technical aspects of the system. A successful pick is not necessarily the end of the perception problem; verifying the part after gripping gives the robot another opportunity to correct positioning before placement.
Reducing Fixed Vision Infrastructure
A conventional bin-picking cell may require fixed cameras positioned above or around the working area. Such installations can introduce additional mechanical mounting, calibration, cabling, and configuration requirements.
With one camera mounted on the robot arm, the same vision system can move between different viewing positions. This can reduce the amount of dedicated vision hardware required for applications involving multiple bins or changing workstation layouts.
The architecture also makes the vision system less dependent on a single fixed viewpoint. This can be useful when the robot needs to inspect a component from different angles during the picking process.
Performance and Production Considerations
According to Inbolt, the system has been deployed in more than five factories and has achieved high uptime and throughput across these production sites. The company reports pick success rates of up to 95% and average processing times below one second per pick.
These figures should be considered application-specific rather than universal performance specifications. Actual results will depend on factors such as component geometry, surface characteristics, bin loading, robot configuration, gripper design, lighting, cycle requirements, and the degree of part occlusion.
For industrial engineers, this distinction is important. Vision performance should ultimately be evaluated against the complete robotic cell rather than against the vision algorithm alone.
AI and 3D Pose Estimation
The system runs on NVIDIA hardware and uses Inbolt's AI-based robot guidance models for real-time pose estimation and trajectory correction.
The technical objective is not simply to recognize an object. The system must determine where the object is, how it is oriented, whether a grasp is physically accessible, and how the robot should move after the part has been picked.
This makes the combination of 3D perception, pose estimation, robot motion, and feedback control more important than image recognition in isolation.
Where This Architecture Can Add Value
The approach is well suited to applications where part presentation is difficult to standardize. Examples include randomly loaded components, bins that can shift position, variable part orientations, and production environments where multiple part geometries must be handled.
Another potential advantage is deployment across multiple stations. If the same robot and vision architecture can be configured for different bins and part families, engineering teams may be able to reduce the amount of dedicated vision infrastructure required for each workstation.
From an automation perspective, the most interesting benefit is therefore not simply faster picking. It is the possibility of reducing the amount of mechanical and programming effort required to maintain performance when the physical environment changes.
Engineering Perspective
The development of on-arm 3D vision reflects a broader shift in industrial robotics: perception is becoming more tightly integrated with robot motion.
Traditional automation often attempts to eliminate variability through fixtures, fixed camera positions, controlled part presentation, and predefined trajectories. AI-based vision takes the opposite approach by allowing the robot to observe more variation and compensate for it through software and feedback.
That does not eliminate the need for sound mechanical design, gripper selection, robot programming, or process engineering. Instead, it changes where some of the complexity is handled. For applications with genuinely variable part presentation, moving part of that complexity from mechanical constraints into real-time perception and motion control can be a practical engineering trade-off.
Conclusion
Inbolt's bin-picking solution demonstrates how on-arm 3D vision and AI can be combined with closed-loop robot control for less structured material-handling applications. Its reported sub-one-second processing time, up-to-95% pick success rate, and deployment across multiple factories indicate a focus on production use rather than laboratory demonstrations.
The key technical concept is the continuous interaction between perception and robot motion: the system observes the available grasp, verifies the component after gripping, and adjusts the placement trajectory as required. This approach can provide a more adaptable alternative to bin-picking architectures that depend heavily on fixed cameras and predefined grasp conditions.
