Gathering a corpus of multimodal computer-mediated meetings with focus on text and audio interaction

Saturnino Luz, Matt Mouley Bouamrane, Masood Masoodian

Research output: Chapter in Book/Report/Conference proceedingConference contributionProfessional

3 Citations (Scopus)

Abstract

In this paper we describe the gathering of a corpus of synchronised speech and text interaction over the network. The data collection scenarios characterise audio meetings with a significant textual component. Unlike existing meeting corpora, the corpus described in this paper emphasises temporal relationships between speech and text media streams. This is achieved through detailed logging and time stamping of text editing operations, actions on shared user interface widgets and gesturing, as well as generation of speech activity profiles. A set of tools has been developed specifically for these purposes which can be used as a data collection platform for the development of meeting browsers. The data gathered to date consists of nearly 30 hours of recorded audio and time stamped editing operations and gestures.

Original languageEnglish
Title of host publicationProceedings of the 5th International Conference on Language Resources and Evaluation, LREC 2006
Pages407-412
Number of pages6
Publication statusPublished - 2006
MoE publication typeD3 Professional conference proceedings
EventInternational Conference on Language Resources and Evaluation - Genoa, Italy
Duration: 22 May 200628 May 2006
Conference number: 5

Conference

ConferenceInternational Conference on Language Resources and Evaluation
Abbreviated titleLREC
Country/TerritoryItaly
CityGenoa
Period22/05/200628/05/2006

Fingerprint

Dive into the research topics of 'Gathering a corpus of multimodal computer-mediated meetings with focus on text and audio interaction'. Together they form a unique fingerprint.

Cite this