RGB-Depth SLAM Review 2018
1、简介
2、TAXONOMY OF DIFFERENT METHODS(不同方法的分类)
Simultaneous Localization and Mapping (SLAM) have made the real time dense reconstruction possible increasing theprospects of navigation, tracking, and augmented reality problems. Some breakthroughs have been achieved in this regard during past few decades and more remarkable works are still going on. This paper presents an overview of SLAM approaches that have been developed till now. Kinect Fusion algorithm, its variants and further developed approaches are discussed in detailed. The algorithms and approaches are compared for their effectiveness in tracking and mapping based on Root Mean Square error over online available datasets.
同时定位与建图(SLAM)使得实时密集重建成为可能,增加了导航、跟踪和增强现实问题的前景。在过去的几十年里,在这方面已经取得了一些突破,更多出色的工作仍在继续。本文综述了近年来发展起来的SLAM方法。详细讨论了Kinect融合算法及其变体和进一步发展的方法,比较了基于均方根误差的在线可用数据集在跟踪和建图的有效性。
2.1 RGB-Depth Mapping
As far as an optimal perception of phenomenal consciousness is concerned, theories based on representation of the mind are based on models of the information processing paradigm [1]. These are as much in correspondence to the neurobiological or functional theories, at this point we are confronted with several arguments on the basis of inversion or absent qualia [2]. Such considerations exhibit a preceding pattern based on the assumption of holding complete knowledge of the neural and functional states that are in subservience to the occurrence of the consciousness that is phenomenal. This can still be conceived as the neural states which are also defined as the states with similar casual responsibilities or with similar representational function [3], [4].
就现象意识的最佳感知而言,基于心智表征的理论是基于信息处理范例[1]的模型。这些都与神经生物学或功能理论相一致,在这一点上,我们面临着几个基于反向或缺少特性[2]的争论。这些考虑显示了一个基于对神经和功能状态的完整知识的假设的模式,这些状态服从于现象性意识的出现。这仍然可以被理解为神经状态,它也被定义为具有相似的偶然职责或相似的表征功能[3],[4]的状态。
2.2 Spatially Extended and Moving Volume Kinetic Fusion(空间扩展和移动体积的动态融合)
These occur with no phenomenal content in any way or such states being accompanied by contents that are phenomenal with broad variation from the usual ones. In definition, visual information processing entails the visual cognitive skills that permit us the processing and interpretation of meaning from visualized information that we attain through eye sight. Therefore, visual perception plays are vital role in aspects of cognitive and intelligence skills such as spelling, math and reading ( [5]). On the other hand, visual perceptual deficits can lead to challenges in learning, recognition and remembrance of letters, wording, and confusion of likeness as well as minor variations in addition to differentiating the main ideal from the details of insignificance.
这些都是在没有现象性内容的情况下发生的,或者这种状态伴随着现象性的内容,这些内容与通常的内容有很大的不同。从定义上说,视觉信息处理需要视觉认知技能,这使得我们能够处理和解释通过视觉获得的视觉化信息。因此,视觉感知在认知和智力技能如拼写、数学和阅读([5])方面起着至关重要的作用。另一方面,视觉感知缺陷会导致学习、识别和记忆字母、措辞、相似的混淆等方面的挑战,除了将主要理想与无关紧要的细节区分开外,还会导致细微的变化。
2.3 Scalable Real-Time Volumetric Reconstruction(可扩展的实时体积重建)
Visual perceptual processing can be sub segmented into the categories that comprise of visual discrimination, figure grounding, closure, memory, sequential memorization, constancy, spatial relations as well as visual motor integration. Note should be taken of perception as active procedures of location and information extraction form the setting while learning entails the procedures of acquisition of information
through experiences of information storage. In which case, thought is the manipulative stance upon information for solving challenges ( [6]). Such that it is eased to extract information (perception) which creates an ease in thought procedures becoming. In overall it is accepted that human vision takes the form of extreme powerful processing of information towards facilitation of the interaction of the world
that surrounds us. However, even in the face of extended and extensive efforts of research encompassing multiple fields of exploration, the fundamentals that underlay as well as operational principles of visual information procedures remain largely unknown.
视觉感知处理可细分为视觉辨别、图形根植、闭合、记忆、顺序记忆、恒常性、空间关系以及视觉运动整合等类别。应注意的是,感知是环境中位置和信息提取的主动过程,而学习是通过信息存储经验获取信息的过程。在这种情况下,思想是解决挑战的信息操纵立场([6])。这样它就能很容易地提取信息(感知),从而在思想过程的变化中变得简单。总的来说,人们普遍认为,人类的视觉是以极其强大的信息处理方式来促进我们周围世界的相互作用。然而,即使面对包括多个探索领域的广泛而广泛的研究努力,视觉信息程序的基本原理和操作原则在很大程度上仍然是未知的。
2.4 Segmentation-Based RGB-D Mapping(基于分割的RGB-D建图)
We are still not able to ascertain the origin and distance along the route from eyes to the sensory input area known as the cortex. It is in this area that the conversion into object meaningful representation is undertaken under conscious manipulation of the brain ( [7]). Nearly half of the human brain in the cerebral cortex region is charged with the processes of visual information although even with extended and extensive research efforts that are encompassed a conundrum still persists. Present theories on visual information processing are held in the consideration of human visual information processing being interplay of the two inversely directed procedural streams.
我们仍然无法确定从眼睛到感觉输入区(即大脑皮层)之间的起点和距离。正是在这个区域,在大脑的有意识控制下([7]),转换成有意义的物体表征。人类大脑皮层近一半的区域负责视觉信息的处理,尽管围绕这一谜题展开了广泛的研究,但这一谜题仍然存在。现有的视觉信息处理理论认为,人的视觉信息处理是两种逆向过程流的相互作用。
2.5 B-D Visual Odometry(B-D视觉里程计)
2.6 Elastic Fusion(弹性融合)
Past research has presented a demonstration of distance and physical enviroment being among the aspects that impairs processing of information, although it remains unknown whether such impairment is on all the levels of information processing or in the onset states instead of the later stages. Those faced with the condition of mapping algorithms suffer from deficiencies of attention that are impairment
to the capability of selective procedures of visual information that is incoming. The early levels of information processing are held in the description of being those that entail the detection as well as response of simplified stimuli. An assignment on the assessment of such function is the inspection time that has previously been demonstrated to entail sensitivity to pharmacological agents.
过去的研究已经表明,距离和物理环境是影响信息处理的因素之一,尽管还不清楚这种损害是在信息处理的所有层次上,还是在开始阶段而不是后期。那些面临建图算法条件的人遭受着注意力缺陷的困扰,这损害了视觉信息传入的选择过程的能力。信息处理的早期阶段是那些需要对简化的刺激进行检测和反应的描述。评估这种功能的一个作业是检查时间,之前已经被证明对药理学药物敏感。
This is as well as being the most reliable and validated within the cultural fairness of information processing measures of cognitive ability ( [5]). Past assessment findings have also presented the impact of nicotine on information procedures as being held in the overall regard in the form of a measure of speed within the early levels of information processing. These include the speed of visual encoding that comprises of the ability of making observations or inspections on sensory input on which the discrimination of relative magnitude rests. This is in contrast to assignments such as reaction time which is summarization entails the involvement of increased response oriented measures of complete decision making time that comprise of total information processing.
这也是在认知能力([5])的信息处理测量的文化公平性中最可靠和最有效的。过去的评估结果还表明,尼古丁对信息处理程序的影响在总体上以信息处理早期阶段的速度衡量标准的形式存在。这包括视觉编码的速度,这种速度包括对感官输入进行观察或检查的能力,而相对大小的辨别就建立在这种能力上。这与诸如反应时间这样的作业形成了对比,反应时间是一种总结,需要增加响应导向的完整决策时间的度量,包括整个信息处理。
Although, there is no research of examination of the impacts administration of 3D scene construction in a similar response, there are limited studies based on the examination of the impacts of 3D scene construction in the early stages of information processing with utilization of other assignments ( [6]). With the application of visual tracking assignments, it was ascertained that the speed of detection experienced impairment from 3D scene construction that that these impacts where greater in dual task settings with comparison to single task settings. Such outcomes have been held in the description of being the deleterious impacts of 3D scene construction on the centralized processing capacity and on information processing availability on the capacity of information processing with time.
虽然目前还没有类似响应下的三维场景建设影响管理的研究,但是利用其他作业([6])对信息处理早期阶段的三维场景建设影响进行检查的研究有限。通过视觉跟踪任务的应用,确定了三维场景构建对检测速度的影响,这些影响在双任务设置时比单任务设置时更大。这些结果被认为是三维场景构建对集中处理能力和信息处理可用性随时间的变化对信息处理能力的有害影响。
Further investigations of early information processing are based on the examination of the mismatched negative component of auditory event relation potential as well as reports of reduced dosage of 3D scene construction attenuation of the event relation potential signal. In this case, the mismatched negative component suppression was solid within stimuli deviation as reduced which the indication of relatively reduced blood 3D scene construction concentration is. The detection of minimal deviations for instance that needed in the course of the inspection time assignment more so in case of hampering in which case similar outcomes have been discovered in simplified reaction time assignments with double level of intensified stimuli. These studies produced outcomes of an increase in response time as well as the impairment of stimuli detection which is a suggestion of the influence on sensory perceptual procedures and the measure of attentiveness ( [7]).
对早期信息处理的进一步研究是基于对听觉事件关系电位负分量不匹配的检验,以及事件关系电位信号的三维场景构建衰减剂量减少的报道。在本例中,不匹配的阴性成分抑制在刺激偏差范围内呈实性降低,提示血三维场景构建浓度相对降低。例如,在检查时间分配的过程中需要检测最小偏差,在检查受阻的情况下更是如此,在这种情况下,在双重强化刺激的简化反应时间分配中发现了类似的结果。这些研究的结果是反应时间的增加以及刺激检测的损伤,这暗示了对感觉知觉过程的影响和注意力的测量([7])。
2.7 Bundle Fusion
Current discoveries in the arena of visual information processing are based on the reflection of the elementary principles of vision as well as the utilization of visual information based on cognitive attributes. This is based on the notion of such work leading to the verge of development based on the grounds of optimism within the several computational theories of sophistication that incorporate data that
is neurobiological and behavioral. These theories entail the flourishing of the skillful exploitation of the neural-imaging and computation of simulative technologies, these permits answering of questions that are subtle regarding the component subsystems within vision.
当前在视觉信息处理领域的发现是基于对视觉基本原理的反映和基于认知属性的视觉信息的利用。这是基于这类工作的概念导致了发展的边缘基于乐观的基础在几个复杂的计算理论中结合了神经生物学和行为学的数据。这些理论带来了对神经成像和模拟技术计算的熟练开发,这些允许回答关于视觉中的组成子系统的微妙问题。
3、MOST RELEVANT METHODS AND DESCRIPTION OF THEIR NOVELTIES AND CONTRIBUTIONS AND WHY ARE THEY PUBLISHED(大多数相关的方法和他们的发现和贡献的描述以及为什么他们被出版)
3.1 GRAPH SLAM
This algorithm applies information matrices sparsely production by the generation of graphs using observed interdependencies in case the observations are connected and if they contain information about the similar landmark. Graph SLAM allows for the capability of constructing a map from an environment while simultaneously creating associated localization with the map for navigation in unknown settings when external referencing systems such as GPS are absent. This intuitive approach utilizes a graph with nodes in correspondence to the robot poses at varied points within time and whose edges are representative of the constraint in between the poses. The latter is gained from environment observations of from movement actions as performed by the robot. Upon construction of a graph, the map could be computed by searching the nodes spatial configuration that is notably consistent with modeled measurements by the edges.
该算法利用观察到的相互依赖关系生成的图来生成稀疏的信息矩阵,如果这些图中包含关于相似地标的信息,则这些图是相互关联的。Graph SLAM允许从环境中构造地图,同时在没有GPS等外部引用系统的情况下创建与地图相关的定位,用于未知设置的导航。这种直观的方法利用了一个图形,其中的节点对应于机器人在不同时间点的姿态,其边缘代表姿态之间的约束。后者是通过对机器人所执行的运动动作的环境观察而获得的。在构建一个图之后,可以通过搜索节点空间配置来计算该地图,该节点空间配置与通过边缘建模的测量值非常一致。
From the image above, we note that particular nodes within the graph are in correspondence to the pose of the robot. Proximal poses are linked by the edges with model spatial constraints between the robot poses that are derived from measurements among the consecutive poses of model odometry measurements. This is whereas the other edges are representative of the spatial constraints based from several observations of the similar section of the environment. The graph-based SLAM method develops a simple estimation challenge by abstraction of raw sensor readings. These readings as substituted by the graph edges which are viewed as ”virtual measurements”. Increased detail within an edge between the two nodes holds the label of a probability distribution over locations that are relative to the two
poses with conditioning to mutual measurements.
从上面的图像中,我们注意到图中的特定节点与机器人的姿态相对应。近端位姿由机器人位姿之间带有模型空间约束的边连接起来,这些边来自于模型测程测量的连续位姿之间的测量值。而其他边缘则是根据对环境相似部分的几次观察而形成的空间约束的代表。基于图的SLAM方法通过提取原始传感器读数,提出了一个简单的估计挑战。这些读数被图形边缘所代替,这些边缘被视为“虚拟测量”。在两个节点之间的边界内增加的细节保持了位置上的概率分布的标签,这些位置是相对于两个位置的,条件是相互测量。
3.2 RGB-D Camera-Based Parallel Tracking and Meshing(基于RGB-D相机的并行跟踪与网格划分)
Visual real-time tracking in regard to established and unknown scenes is critical as well as an incontrovertible aspect in vision-based AR applications. Multiple algorithm contributions over the years. It is at this point that we introduce RGB-D Camera-Based Parallel Tracking and Meshing as an adaptation and updating of the algorithms utilized in estimating the motion of the camera as well as AR in accordance to the availability of the end user in computational abilities in permitting to gain impressive tracking outcomes in limited AR workspaces. The fact is that estimation of camera motion using environment tracking as well as parallel constructing feature based sparse mapping that creates a possibility in part to the generalization of multi-core processors found in desktop and laptop computers. Of recent is has been revealed that increased computation power within a singular standard of a hand-held video camera is connected to a powerful computer using computational power gained from the Graphics Processing Unit (GPU). The possibility to attain a dense representation of a desktop setting as well as increased texturing scenery whereas as undertaking tracking with the use RGB-D Camera-Based Parallel Tracking and Meshing. The online created map density can be increased with the use stereo-dense matching in addition to GPU founded implementations as shown by GPU to be utilized for effective replacement of the global bundle adjustment aspects of SLAM optimized based systems for instance RGBD Camera-Based Parallel Tracking and Meshing as well as inherent parallelization refinement with step founded Monte Carlo simulations therefore freeing tools on the CPU for other assignments.
在基于视觉的增强现实应用中,对已建立的和未知的场景进行实时跟踪是关键的,也是一个不容置疑的方面。多年来的多重算法贡献。在这一点上,我们介绍基于RGB-D相机平行跟踪和啮合的适应和更新算法用于估计摄像机的运动以及AR按照最终用户的可用性计算能力允许在有限的基于“增大化现实”技术获得令人印象深刻的跟踪结果工作区。事实上,使用环境跟踪和基于并行构造特征的稀疏映射来估计摄像机的运动,这在一定程度上创造了在台式机和笔记本电脑中发现的多核处理器泛化的可能性。最近的一项研究表明,在一个单一标准的手持相机内增加的计算能力与一台强大的计算机相连接,该计算机使用从图形处理单元(GPU)获得的计算能力。实现一个密集的桌面设置和增加纹理风景的可能性,而跟踪使用基于RGB-D相机的并行跟踪和网格。创建的在线地图密度可以增加使用stereo-dense匹配除了GPU实现如图所示由GPU用于有效替代的全球大满贯的束调整方面基于优化的系统例如基于摄像头RGBD平行跟踪和啮合以及固有的并行细化步骤建立蒙特卡洛模拟因此释放工具在CPU上的其他作业。