Current open-vocabulary object detection performs well in clean scenarios. However, in actual use cases, most images are from suboptimal viewpoints, contain many occlusions, or do not contain relevant information at all. Even when aggregating information from multiple views, current state-of-the-art object detection methods consistently suffer from relatively obvious false positives. We explore the use of 3D-predictive tools to address the limitations of current object detection algorithms.