Relation-prediction-from-embeddings
2022Description
Relation-prediction-from-embeddings
Relation-prediction-from-embeddings
A model for Unsupervised Disentanglement of Syntax and Semantics (code)
This repo contains the code for our paper Exploiting Inductive Bias in Transformers for Unsupervised Disentanglement of Syntax and Semantics with VAEs
Prédiction de sémantiques sur Wordnet
Named Entity Recognition (NER) is an important task in Natural Language Processing with applications in many domains. We present an algorithm that efficiently finds the set of non-overlapping spans that maximizes a global score, given a list of candidate span representations that avoids span overlap.
MBart-based model for the indexing of scientific documents. Presented to NLDB 2022. "Transformer-based models for the Automatic Indexing of Scientific Documents in French" José Angel Gonzalez, Davide Buscaldi, Lluis Hurtado and Emilio Sanchis
Un analyseur en dépendances d'ordre supérieur fondé sur la descente par coordonnées
Named Entity Recognition (NER) is an important task in Natural Language Processing with applications in many domains. We present a named entity recognition system in which we output a set of spans (i.e., segmentations) by maximizing a global score. During training, we optimize our model by maximizing the probability of the gold segmentation. During inference, we use dynamic programming to select the best segmentation under a linear time complexity.
Reducing Overlapping in Span-based NER Using Graph Neural Networks (code)
Code for Reducing Overlapping in Span-based NER Using Graph Neural Networks
Dynamic Query Representation for Extractive Question Answering (code)
Extractive question answering (ExQA) is an essential task for Natural Language Processing. To address this task, we propose DyREx, a generalization of the vanilla approach where we dynamically compute query vectors given the input, using an attention mechanism through transformer layers.
A Large-Scale Knowledge Graph of Research Entities and Claims in Computer Science
Un analyseur en dépendances d'ordre supérieur fondé sur un réseau d'inférence.
a Pretrained Arabic Sequence-to-Sequence Model for Abstractive Summarization (code)
averaging and clustering effects in mixture of experts for graph-based dependency parsing (code)
Mixture of Experts, training for mixture with clustering. Un VAE texte-vers-texte pour la génération contrôlée.
Using GCN to build a vectorial representation of the sentence, starting from its dependency graph. This representation can be used to carry out other tasks. Work is still in progress, the model was used for the PreTens22 challenge
La recherche d'experts parmi les migrations hautement qualifiées est un enjeu crucial pour les pays en développement. Cet système implémente une méthode d'apprentissage par renforcement profond pour répondre à cette problématique à partir de résultats des moteurs de recherche sur le web. Les résultats obtenus pourront être utilisés par des sociologues de la migration pour mieux comprendre les diasporas des savoirs, ainsi que par le pays en développement pour localiser ses experts formés à l'étranger. Notre système effectue des requêtes envers des moteurs de recherche pour y extraire les informations concernant l'institution et l'année d'affiliation de chaque expert. L'objectif de ce travail est de définir un navigateur intelligent capable d'assister cette recherche en générant et en observant le moins possible de requêtes automatiques. Nous utilisons comme navigateur un Deep-Q Network avec deux architectures basées sur des réseaux de neurones pour approximer la valeur de la Q-value fonction.
General-domain Summary Evaluation by Relevance Analysis (code)
We present GeSERA, an open-source improved version of SERA for evaluating automatic extractive and abstractive summaries from the general domain. SERA is based on a search engine that compares candidate and reference summaries (called queries) against an information retrieval document base (called index). SERA was originally designed for the biomedical domain only, where it showed a better correlation with manual methods than the widely used lexical-based ROUGE method. In this paper, we take out SERA from the biomedical domain to the general one by adapting its content-based method to successfully evaluate summaries from the general domain. First, we improve the query reformulation strategy with POS Tags analysis of general-domain corpora. Second, we replace the biomedical index used in SERA with two article collections from AQUAINT-2 and Wikipedia.
plate-forme d'intégration d'outils de traitement automatique des langues
La plateforme ChêneTAL a été conçue pour permettre la mise en place de chaînes hétérogènes de Traitement Automatique des Langues (TAL) en intégrant des logiciels existants en gestion et manipulation de corpus avec des modèles plus récents d’Intelligence Artificielle (IA) et en proposant une interface simple et intuitive pour les chercheur·euse·s de la communauté en Traitement Automatique des Langues (TAL) et pour ceux en Linguistique/Sciences Humaines et Sociales non spécialistes en informatique. Code: https://depot.lipn.univ-paris13.fr/garciaflores/ch-netal
Exploiting Complementarities of Different Dependency Representations (code)
validation app of comparable corpus for the Maya Yucatec language (code)
The objective of this project is to develop resources for the automatic translation of the Maya Yucatec language.
Traducteur espagnol - Wixárika (langue amérindienne) (démo)
Traducteur espagnol - Wixárika (langue amérindienne).
La plateforme Néoveille a pour objectif d'offrir un outil de détection et de suivi des néologismes dans la presse en ligne et plus généralement l'ensemble des données disponibles sur le web. Le projet a été financé pour trois ans (juin 2015 - juin 2018) par la COMUE Sorbonne Paris Cité (regroupant plusieurs laboratoires de Sorbonne-Paris-Cité (LIPN, LDI, CLILLAC-ARP, ERTIM), les acteurs du groupe EMPNEO et l'Université de São Paulo (USP)), puis financé par la Direction Générale à la Langue Française et aux Langues de France (DGLF-LF). Le projet propose : Une interface de gestion de sources de presse en ligne (format RSS) : les sources sont ensuite récupérées une fois par jour et les néologismes automatiquement détectés ; Une interface de validation/invalidation des néologismes détectés automatiquement dans la phase précédente ; Une interface de suivi des néologismes validés, avec une visualisation des contextes et un suivi par différents indicateurs métalinguistiques (pays, journal, domaine);
Character-based Models for the Automatic Misogyny Identification Task (code)
tag prediction and semantic indexing system (code)
Document classification is often meant to serve as semantic indexing to help readers finding documents related to a given topic. However, the quality of indexing typically deteriorates with time: some categories are misused or forgotten by indexers, others become obsolete or too general to be useful. We implemented a semantic indexing system as an algorithm that guides indexers in restructuring their indexes. Focus is put on the reader’s rather than on the annotator’s point of view.
Topic mining data processing and visualization (mostly on wikipedia)
Topic mining data processing and visualization (mostly on wikipedia)
Robot Experience Stories Generator (code)
The aim of this system is to provide service robots with natural language capabilities to produce a Robot Experience Story for its human interlocutors. Golfred stories are narratives composed of the robot's holistic perception of a recently performed task: navigation, visual perceptions and action descriptions. We implemented with a narrate dialog model specifying the composition of situations necessary for a service robot to transform its task history record into a narrative knowledge representation. We provide SitLog algorithms allowing to analyze the robot's situation and behaviors sequence in order to generate a Golfred story of the task. Both the dialogue model and the algorithms can be embedded as compositional behaviors in any other SitLog task structure. We instantiated our model into the Golem service robot framework.
plateforme d'aide à la construction de ressources termino-ontologiques à partir de ressources textuelles. (http://lipn.fr/terminae/index.php/Main_Page)
un environnement pédagogique pour la simulation de réseaux locaux
Démonstration à la conférence EIAH 2007 (Environnements Informatiques pour l'Apprentissage Humain) Projet E-Learning Université Paris 13