CINECA IRIS Institutional Research Information System

The process of extracting relevant technical information from patents or technical literature is as valuable as it is challenging. It deals with highly relevant information extraction from a corpus of documents with particular structure, and a mix of technical and legal jargon. Patents are the wider free source of technical information where homogeneous entities can be found. From a technical perspective the approaches refer to Named Entity Recognition (NER) and make use of Machine Learning techniques for Natural Language Processing (NLP). However, due to the large amount of data, to the complexity of the lexicon, the peculiarity of the structure and the scarcity of the examples to be used to feed the machine learning system, new approaches should be studied. NER methods are increasing their performances in many contexts, but a gap still exists when dealing with technical documentation. The aim of this work is to create an automatic training sets for NER systems by exploiting the nature and structure of patents, an open and massive source of technical documentation. In particular, we focus on collecting the context where users of the invention appear within patents. We then measure to which extent we achieve our goal and discuss how much our method is generalizable to other entities and documents.

A simple and fast method for Named Entity context extraction from patents

Puccetti G.^Primo;Chiarello F.^Secondo;Fantoni G.^Ultimo

2021-01-01

Abstract

The process of extracting relevant technical information from patents or technical literature is as valuable as it is challenging. It deals with highly relevant information extraction from a corpus of documents with particular structure, and a mix of technical and legal jargon. Patents are the wider free source of technical information where homogeneous entities can be found. From a technical perspective the approaches refer to Named Entity Recognition (NER) and make use of Machine Learning techniques for Natural Language Processing (NLP). However, due to the large amount of data, to the complexity of the lexicon, the peculiarity of the structure and the scarcity of the examples to be used to feed the machine learning system, new approaches should be studied. NER methods are increasing their performances in many contexts, but a gap still exists when dealing with technical documentation. The aim of this work is to create an automatic training sets for NER systems by exploiting the nature and structure of patents, an open and massive source of technical documentation. In particular, we focus on collecting the context where users of the invention appear within patents. We then measure to which extent we achieve our goal and discuss how much our method is generalizable to other entities and documents.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno
	
				2021
			
	Codice DOI
	
				https://dx.doi.org/10.1016/j.eswa.2021.115570
			
	Tutti gli autori
	
						Puccetti, G.; Chiarello, F.; Fantoni, G.
					
	Appare nelle tipologie:
	
				1.1 Articolo in rivista

File in questo prodotto:

File	Dimensione	Formato
A Simple and Fast Method for Named Entity Context Extraction from Patents.pdf non disponibili Descrizione: Acceptance letter Tipologia: Altro materiale allegato Licenza: NON PUBBLICO - accesso privato/ristretto Dimensione 37.58 kB Formato Adobe PDF Visualizza/Apri Richiedi una copia	37.58 kB	Adobe PDF	Visualizza/Apri Richiedi una copia
A_simple_and_fast_method_AAM.pdf accesso aperto Tipologia: Documento in Post-print Licenza: Creative commons Dimensione 494.74 kB Formato Adobe PDF Visualizza/Apri	494.74 kB	Adobe PDF	Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11568/1107028

Citazioni

ND

23

19

social impact