Skip to content

Equipe RCLN

Logiciels


logiciels

QKVAE

2022

A model for Unsupervised Disentanglement of Syntax and Semantics (code)

Ghazi Felhi, Joseph Le Roux, Djamé Seddah

Description

This repo contains the code for our paper Exploiting Inductive Bias in Transformers for Unsupervised Disentanglement of Syntax and Semantics with VAEs

MBartCopyGenerator

2022

Jose Angel Gonzalez, Davide Buscaldi, Lluis Hurtado, Emilio Sanchis

Description

MBart-based model for the indexing of scientific documents. Presented to NLDB 2022. "Transformer-based models for the Automatic Indexing of Scientific Documents in French" José Angel Gonzalez, Davide Buscaldi, Lluis Hurtado and Emilio Sanchis

Global Span Selection for Named Entity Recognition (code)

2022

Urchade Zaratiana, Niama Elkhbir, Pierre Holat, Nadi Tomeh, Thierry Charnois

Description

Named Entity Recognition (NER) is an important task in Natural Language Processing with applications in many domains. We present a named entity recognition system in which we output a set of spans (i.e., segmentations) by maximizing a global score. During training, we optimize our model by maximizing the probability of the gold segmentation. During inference, we use dynamic programming to select the best segmentation under a linear time complexity.

GNNer

2022

Reducing Overlapping in Span-based NER Using Graph Neural Networks (code)

Urchade Zaratiana, Nadi Tomeh, Pierre Holat, Thierry Charnois

Description

Code for Reducing Overlapping in Span-based NER Using Graph Neural Networks

DyREx

2022

Dynamic Query Representation for Extractive Question Answering (code)

Urchade Zaratiana, Niama Elkhbir, Dennis Hernando Nunez Fernandez, Pierre Holat, Nadi Tomeh, Thierry Charnois

Description

Extractive question answering (ExQA) is an essential task for Natural Language Processing. To address this task, we propose DyREx, a generalization of the vanilla approach where we dynamically compute query vectors given the input, using an attention mechanism through transformer layers.

CS-KG

2022

A Large-Scale Knowledge Graph of Research Entities and Claims in Computer Science

Davide Buscaldi, Danilo Dessì, Diego Reforgiato Recupero, Francesco Osborne, Enrico Motta

Strength in numbers

2021

averaging and clustering effects in mixture of experts for graph-based dependency parsing (code)

Xudong Zhang, Joseph Le Roux, Thierry Charnois

Description

Mixture of Experts, training for mixture with clustering. Un VAE texte-vers-texte pour la génération contrôlée.

RENFO / DiasporasRL (code)

2021

Jessica Lopez Espejel, Pegah Alizadeh, Jorge Garcia Flores

Description

La recherche d'experts parmi les migrations hautement qualifiées est un enjeu crucial pour les pays en développement. Cet système implémente une méthode d'apprentissage par renforcement profond pour répondre à cette problématique à partir de résultats des moteurs de recherche sur le web. Les résultats obtenus pourront être utilisés par des sociologues de la migration pour mieux comprendre les diasporas des savoirs, ainsi que par le pays en développement pour localiser ses experts formés à l'étranger. Notre système effectue des requêtes envers des moteurs de recherche pour y extraire les informations concernant l'institution et l'année d'affiliation de chaque expert. L'objectif de ce travail est de définir un navigateur intelligent capable d'assister cette recherche en générant et en observant le moins possible de requêtes automatiques. Nous utilisons comme navigateur un Deep-Q Network avec deux architectures basées sur des réseaux de neurones pour approximer la valeur de la Q-value fonction.

GeSERA

2021

General-domain Summary Evaluation by Relevance Analysis (code)

Jessica Lopez Espejel, Gaël De Chalendar, Jorge Garcia Flores, Thierry Charnois, Ivan Vladimir Meza Ruiz

Description

We present GeSERA, an open-source improved version of SERA for evaluating automatic extractive and abstractive summaries from the general domain. SERA is based on a search engine that compares candidate and reference summaries (called queries) against an information retrieval document base (called index). SERA was originally designed for the biomedical domain only, where it showed a better correlation with manual methods than the widely used lexical-based ROUGE method. In this paper, we take out SERA from the biomedical domain to the general one by adapting its content-based method to successfully evaluate summaries from the general domain. First, we improve the query reformulation strategy with POS Tags analysis of general-domain corpora. Second, we replace the biomedical index used in SERA with two article collections from AQUAINT-2 and Wikipedia.

ChêneTAL

2021

plate-forme d'intégration d'outils de traitement automatique des langues

Aude Grezka, Jorge Garcia Flores, Thierry Charnois

Description

La plateforme ChêneTAL a été conçue pour permettre la mise en place de chaînes hétérogènes de Traitement Automatique des Langues (TAL) en intégrant des logiciels existants en gestion et manipulation de corpus avec des modèles plus récents d’Intelligence Artificielle (IA) et en proposant une interface simple et intuitive pour les chercheur·euse·s de la communauté en Traitement Automatique des Langues (TAL) et pour ceux en Linguistique/Sciences Humaines et Sociales non spécialistes en informatique. Code: https://depot.lipn.univ-paris13.fr/garciaflores/ch-netal

YuCorrectMaya

2019

validation app of comparable corpus for the Maya Yucatec language (code)

Heba Kaddouh, Maroi Labiodh, Nouha Ghourabi, Jorge Garcia Flores, Alik Hafsa, Cylia Ourtirane

Description

The objective of this project is to develop resources for the automatic translation of the Maya Yucatec language.

Pamparios

2019

Traducteur espagnol - Wixárika (langue amérindienne) (démo)

Jorge Garcia Flores, Hugo Ferreira, Aziz Okotan, Fernando Mantilla, Fayaz Abdoulvahide

Description

Traducteur espagnol - Wixárika (langue amérindienne).

Neoveille, plateforme de repérage et de suivi des néologismes en corpus dynamique (outil en ligne)

2019

Emmanuel Cartier

Description

La plateforme Néoveille a pour objectif d'offrir un outil de détection et de suivi des néologismes dans la presse en ligne et plus généralement l'ensemble des données disponibles sur le web. Le projet a été financé pour trois ans (juin 2015 - juin 2018) par la COMUE Sorbonne Paris Cité (regroupant plusieurs laboratoires de Sorbonne-Paris-Cité (LIPN, LDI, CLILLAC-ARP, ERTIM), les acteurs du groupe EMPNEO et l'Université de São Paulo (USP)), puis financé par la Direction Générale à la Langue Française et aux Langues de France (DGLF-LF). Le projet propose : Une interface de gestion de sources de presse en ligne (format RSS) : les sources sont ensuite récupérées une fois par jour et les néologismes automatiquement détectés ; Une interface de validation/invalidation des néologismes détectés automatiquement dans la phase précédente ; Une interface de suivi des néologismes validés, avec une visualisation des contextes et un suivi par différents indicateurs métalinguistiques (pays, journal, domaine);

Kilroy

2018

tag prediction and semantic indexing system (code)

Ivan Garrido Marquez, Jorge Garcia Flores, François Lévy, Adeline Nazarenko

Description

Document classification is often meant to serve as semantic indexing to help readers finding documents related to a given topic. However, the quality of indexing typically deteriorates with time: some categories are misused or forgotten by indexers, others become obsolete or too general to be useful. We implemented a semantic indexing system as an algorithm that guides indexers in restructuring their indexes. Focus is put on the reader’s rather than on the annotator’s point of view.

min-hashing

2017

Topic mining data processing and visualization (mostly on wikipedia)

Ivan Vladimir Meza Ruiz, Gibran Fuentes-Pineda, Jorge Garcia Flores, Mohamed Chabouni, Zakaria Khezane

Description

Topic mining data processing and visualization (mostly on wikipedia)

Golfred

2017

Robot Experience Stories Generator (code)

Jorge Garcia Flores, Ivan Vladimir Meza Ruiz, Luis Alberto Pineda Cortes

Description

The aim of this system is to provide service robots with natural language capabilities to produce a Robot Experience Story for its human interlocutors. Golfred stories are narratives composed of the robot's holistic perception of a recently performed task: navigation, visual perceptions and action descriptions. We implemented with a narrate dialog model specifying the composition of situations necessary for a service robot to transform its task history record into a narrative knowledge representation. We provide SitLog algorithms allowing to analyze the robot's situation and behaviors sequence in order to generate a Golfred story of the task. Both the dialogue model and the algorithms can be embedded as compositional behaviors in any other SitLog task structure. We instantiated our model into the Golem service robot framework.

Logiciel Terminae - Version 2012

2012

Sylvie Szulman

Description

plateforme d'aide à la construction de ressources termino-ontologiques à partir de ressources textuelles. (http://lipn.fr/terminae/index.php/Main_Page)

Marionnet/Iutoppix

2007

un environnement pédagogique pour la simulation de réseaux locaux

Jean-Vincent Loddo, Thierry Hamon

Description

Démonstration à la conférence EIAH 2007 (Environnements Informatiques pour l'Apprentissage Humain) Projet E-Learning Université Paris 13