
Reading a restaurant menu, scanning a printed document, or extracting text from a laptop screen are becoming increasingly common use cases for AI-powered smartglasses. As these devices evolve, users expect AI assistants to recognize and interpret text accurately, regardless of the environment.
While much of the attention is focused on AI models, text recognition performance depends on the entire vision pipeline. Camera hardware, image processing, software optimization, and the AI model all contribute to how effectively a device can perceive and interpret visual information. Depending on the product, the AI may analyze a still image or process a live camera stream, but in every case the overall user experience results from the interaction of these components rather than the AI model alone.
Evaluating AI Text Recognition in Realistic Conditions
Assessing AI-powered text recognition therefore requires evaluating the complete system rather than the AI model in isolation.
With this objective in mind, DXOMARK developed a benchmark designed to assess text recognition in realistic but controlled conditions. The protocol evaluates the complete vision pipeline from the visual input received by the device through image processing and AI analysis to the final response generated by the smartglasses.
The benchmark covers three common situations that users are likely to encounter in everyday life:
-
- Reading printed text
- Reading and understanding a restaurant menu
- Reading text displayed on a laptop screen


Each use case introduces different challenges. Reading printed text primarily measures how accurately the device recognizes characters of different sizes and styles. A restaurant menu goes one step further by evaluating whether the device can correctly identify dishes, descriptions, and prices without confusing or omitting information. Reading text on a laptop screen presents another challenge, as the glasses must deal with different display brightness levels and contrast.
This article focuses on the first use case, reading printed text, which provides a straightforward way to compare how different smartglasses capture and recognize written information.
To ensure fair and repeatable comparisons, every device was evaluated under controlled conditions. Tests were carried out in DXOMARK’s laboratory using the same printed test chart, the same viewing distance of 40 cm, and three controlled lighting environments representative of everyday use:
-
- Bright daylight (1000 lux)
- Typical indoor lighting (300 lux)
- Low light (20 lux)
Each device also received the exact same prompt:
“Read all the text on the page in front of me exactly as it appears. Do not add, correct, or guess any information.”
The evaluation included six commercially available smartglasses:
Meta Ray-Ban Display, Meta Ray-Ban Gen. 2, Rokid AI Glasses, HTC Vive Eagle, Rayneo X3 Pro, and Quark AI S1.

Several of these devices support multiple AI assistants. To ensure consistency throughout the evaluation, one AI model was selected for each product and used for every test. This approach makes it easier to compare complete smartglasses systems in a repeatable way while minimizing variables that could influence the results.
The benchmark also looks beyond simple recognition accuracy. It evaluates factors such as recognition completeness, response time, prompt understanding, and AI hallucinations. Together, these measurements provide a broader view of how effectively smartglasses can recognize and return textual information in real-world situations.

What the Benchmark Revealed
The benchmark confirmed several trends across the tested devices. Lighting conditions remain an important factor. Recognition performance generally decreased in low light, while larger text was consistently easier to recognize than smaller fonts. Text contrast also influenced performance, with black text on a white background generally producing better results on paper.
The most interesting observation was that devices using the same underlying AI model did not necessarily produce the same results. This indicates that the AI model alone does not determine the overall experience. The quality of image capture and the way visual information is processed before reaching the AI also play a significant role.

Among the evaluated products, the Rokid AI Glasses achieved the highest scores for printed text recognition across all lighting conditions.
In this context, the Rokid AI glasses achieved the highest scores for printed text recognition, regardless of lighting conditions. This advantage could be attributed to Rokid’s over-sharpening strategy, which appears to be applied consistently to both photography and AI use cases. This strategy makes the AI model’s work easier for OCR, but also results in unnatural-looking images for photography use cases as observed in our benchmark.
This suggests that two different image pipelines with different tuning are required to offer a better experience for both photography and AI reading use-case. Since the accuracy of optical character recognition (OCR) depends heavily on the quality of the source image, Rokid’s tuning strategy gives its glasses an advantage in terms of reading performance, even when using an AI model identical to that of its competitors.
