Modelo de extracción de información clave para el Documento Nacional de Identidad Peruano mediante YOLO y TrOCR

dc.contributor.advisorHuanca Torres, Fredy Abel
dc.contributor.authorQuispe Valero, Jhoselyn Noemi
dc.date.accessioned2026-08-31T17:40:55Z
dc.date.embargoEnd2028-08-11
dc.date.issued2026-08-11
dc.description.abstractAutomatic information extraction from identity documents remains a challenging task due to variations in image quality, document layout, and text appearance. This study proposes a key information extraction system for the Peruvian National Identity Document (DNI) by integrating object detection and Transformer-based text inference into a unified processing pipeline. The proposed approach combines YOLO11 for the detection and segmentation of regions of interest (ROI) and TrOCR for inferring the textual content of each identified field. The system was trained and evaluated using multiple document fields, including document number, given names, surnames, date of birth, and sex. Experimental results indicate consistent performance in both detection and text inference tasks, even without incorporating explicit preprocessing or post-processing stages. In particular, the fine-tuning process applied to TrOCR improved the recognition of compound names and Spanish-specific characters, including accented vowels and the letter ñ. The achieved performance was comparable to and, in some cases, surpassed that reported in related studies under equivalent evaluation metrics. Furthermore, the proposed methodology was integrated into a functional web-based application, enabling validation of the complete extraction workflow and the structured organization of the inferred information. Overall, the findings support the feasibility of the proposed system for automatic information extraction from identity documents under real-world conditions.
dc.description.escuelaEscuela Profesional de Ingeniería de Sistemas
dc.description.lineadeinvestigacionGestión de TI
dc.description.sedeJuliaca
dc.formatapplication/pdf
dc.identifier.urihttps://hdl.handle.net/20.500.12840/10609
dc.language.isoeng
dc.publisherUniversidad Peruana Unión
dc.publisher.countryPE
dc.rightsinfo:eu-repo/semantics/embargoedAccess
dc.rights.urihttps://creativecommons.org/licenses/by-nc/4.0/
dc.subjectInformation extraction
dc.subjectIdentity document processing
dc.subjectRegion of interest
dc.subjectObject detection
dc.subjectText inference
dc.subjectTransformer-based models
dc.subject.ocdehttps://purl.org/pe-repo/ocde/ford#1.02.01
dc.titleModelo de extracción de información clave para el Documento Nacional de Identidad Peruano mediante YOLO y TrOCR
dc.typeinfo:eu-repo/semantics/bachelorThesis
renati.advisor.dni01345134
renati.advisor.orcidhttps://orcid.org/0000-0001-7645-7144
renati.author.dni76286033
renati.discipline61200093
renati.jurorTocto Cano, Esteban
renati.jurorMamani Pari, David
renati.jurorHumpiri Flores, Milton Edward
renati.levelhttps://purl.org/pe-repo/renati/level#tituloProfesional
renati.typehttps://purl.org/pe-repo/renati/type#tesis
thesis.degree.disciplineINGENIERÍA DE SISTEMAS
thesis.degree.grantorUniversidad Peruana Unión. Facultad de Ingeniería y Arquitectura
thesis.degree.nameIngeniero de Sistemas

Archivos

Bloque original

Mostrando 1 - 3 de 3
Cargando...
Miniatura
Nombre:
Jhoselyn_Tesis_Licenciatura_2026.pdf
Tamaño:
350,8 KB
Formato:
Adobe Portable Document Format
Cargando...
Miniatura
Nombre:
Autorizacion_de_Publicación.pdf
Tamaño:
129,66 KB
Formato:
Adobe Portable Document Format
Cargando...
Miniatura
Nombre:
Reporte_de_Similitud.pdf
Tamaño:
3,18 MB
Formato:
Adobe Portable Document Format