A Study on Dialog Act Recognition using Character-Level Tokenization

Ribeiro, Eugénio; Ribeiro, Ricardo; de Matos, David Martins

Computer Science > Computation and Language

arXiv:1805.07231 (cs)

[Submitted on 18 May 2018 (v1), last revised 23 Jul 2018 (this version, v2)]

Title:A Study on Dialog Act Recognition using Character-Level Tokenization

Authors:Eugénio Ribeiro, Ricardo Ribeiro, David Martins de Matos

View PDF

Abstract:Dialog act recognition is an important step for dialog systems since it reveals the intention behind the uttered words. Most approaches on the task use word-level tokenization. In contrast, this paper explores the use of character-level tokenization. This is relevant since there is information at the sub-word level that is related to the function of the words and, thus, their intention. We also explore the use of different context windows around each token, which are able to capture important elements, such as affixes. Furthermore, we assess the importance of punctuation and capitalization. We performed experiments on both the Switchboard Dialog Act Corpus and the DIHANA Corpus. In both cases, the experiments not only show that character-level tokenization leads to better performance than the typical word-level approaches, but also that both approaches are able to capture complementary information. Thus, the best results are achieved by combining tokenization at both levels.

Comments:	11 pages, 2 figures, 4 tables, AIMSA 2018
Subjects:	Computation and Language (cs.CL)
ACM classes:	H.1.2; H.3.1; I.2.7
Cite as:	arXiv:1805.07231 [cs.CL]
	(or arXiv:1805.07231v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1805.07231

Submission history

From: Eugénio Ribeiro [view email]
[v1] Fri, 18 May 2018 14:17:07 UTC (91 KB)
[v2] Mon, 23 Jul 2018 13:28:56 UTC (155 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2018-05

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Eugénio Ribeiro
Ricardo Ribeiro
David Martins de Matos

export BibTeX citation

Computer Science > Computation and Language

Title:A Study on Dialog Act Recognition using Character-Level Tokenization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:A Study on Dialog Act Recognition using Character-Level Tokenization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators