Textual Attribute Recognition for Documents


Jyothi Swaroopa Jinka

Abstract

Understanding document content requires not only recognizing textual characters but also capturing textual attributes such as bold, italic, underline, and strikeout, which convey emphasis, hierarchy, and semantic intent. These attributes play a critical role in real-world documents, including legal records, administrative forms, educational materials, and historical manuscripts, where formatting often encodes meaning that cannot be inferred from text alone. Despite their importance, Textual Attribute Recognition (TAR) has received limited attention compared to traditional Optical Character Recognition (OCR), and existing approaches remain inadequate for handling the complexity of real-world, multilingual, and layout-rich documents.

A key limitation of prior TAR methods lies in their treatment of words as isolated visual units. At- tribute recognition, however, is inherently context-dependent, as distinctions such as bold versus regular text or underline versus structural lines often require comparison with neighboring words or surround- ing layout elements. Methods that ignore this context suffer from ambiguity and reduced accuracy. Conversely, approaches that attempt to incorporate full-document context using transformer-based ar- chitectures face significant computational challenges, due to the quadratic complexity of attention over large numbers of word instances. These issues are further exacerbated in multilingual settings, where script diversity, font variations, and document noise introduce additional variability.

 

Year of completion:  May 2026
 Advisor :

Ravi Kiran Sarvadevabhatla


Related Publications


    Downloads

    thesis