Multimedia Topic Models Considering Burstiness of Local Features

Yang XIE; Koji EGUCHI

doi:10.1587/transinf.E97.D.714

IEICE TRANSACTIONS on Information

Open Access
Multimedia Topic Models Considering Burstiness of Local Features

Yang XIE, Koji EGUCHI

Full Text Views

30

Cite this

Free PDF (737.7KB)

Summary :

A number of studies have been conducted on topic modeling for various types of data, including text and image data. We focus particularly on the burstiness of the local features in modeling topics within video data in this paper. Burstiness is a phenomenon that is often discussed for text data. The idea is that if a word is used once in a document, it is more likely to be used again within the document. It is also observed in video data; for example, an object or visual word in video data is more likely to appear repeatedly within the same video data. Based on the idea mentioned above, we propose a new topic model, the Correspondence Dirichlet Compound Multinomial LDA (Corr-DCMLDA), which takes into account the burstiness of the local features in video data. The unknown parameters and latent variables in the model are estimated by conducting a collapsed Gibbs sampling and the hyperparameters are estimated by focusing on the fixed-point iterations. We demonstrate through experimentation on the genre classification of social video data that our model works more effectively than several baselines.

Publication: IEICE TRANSACTIONS on Information Vol.E97-D No.4 pp.714-720

Publication Date: 2014/04/01

Publicized

Online ISSN: 1745-1361

DOI: 10.1587/transinf.E97.D.714

Type of Manuscript: Special Section PAPER (Special Section on Data Engineering and Information Management)

Category

Authors

Yang XIE
Kobe University
Koji EGUCHI
Kobe University

Keyword

topic models, multimedia, word burstiness, Dirichlet compound multinomials

Cite this

Copy

Yang XIE, Koji EGUCHI, "Multimedia Topic Models Considering Burstiness of Local Features" in IEICE TRANSACTIONS on Information, vol. E97-D, no. 4, pp. 714-720, April 2014, doi: 10.1587/transinf.E97.D.714.
Abstract: A number of studies have been conducted on topic modeling for various types of data, including text and image data. We focus particularly on the burstiness of the local features in modeling topics within video data in this paper. Burstiness is a phenomenon that is often discussed for text data. The idea is that if a word is used once in a document, it is more likely to be used again within the document. It is also observed in video data; for example, an object or visual word in video data is more likely to appear repeatedly within the same video data. Based on the idea mentioned above, we propose a new topic model, the Correspondence Dirichlet Compound Multinomial LDA (Corr-DCMLDA), which takes into account the burstiness of the local features in video data. The unknown parameters and latent variables in the model are estimated by conducting a collapsed Gibbs sampling and the hyperparameters are estimated by focusing on the fixed-point iterations. We demonstrate through experimentation on the genre classification of social video data that our model works more effectively than several baselines.
URL: https://global.ieice.org/en_transactions/information/10.1587/transinf.E97.D.714/_p

Copy

@ARTICLE{e97-d_4_714,
author={Yang XIE, Koji EGUCHI, },
journal={IEICE TRANSACTIONS on Information},
title={Multimedia Topic Models Considering Burstiness of Local Features},
year={2014},
volume={E97-D},
number={4},
pages={714-720},
abstract={A number of studies have been conducted on topic modeling for various types of data, including text and image data. We focus particularly on the burstiness of the local features in modeling topics within video data in this paper. Burstiness is a phenomenon that is often discussed for text data. The idea is that if a word is used once in a document, it is more likely to be used again within the document. It is also observed in video data; for example, an object or visual word in video data is more likely to appear repeatedly within the same video data. Based on the idea mentioned above, we propose a new topic model, the Correspondence Dirichlet Compound Multinomial LDA (Corr-DCMLDA), which takes into account the burstiness of the local features in video data. The unknown parameters and latent variables in the model are estimated by conducting a collapsed Gibbs sampling and the hyperparameters are estimated by focusing on the fixed-point iterations. We demonstrate through experimentation on the genre classification of social video data that our model works more effectively than several baselines.},
keywords={},
doi={10.1587/transinf.E97.D.714},
ISSN={1745-1361},
month={April},}

Copy

TY - JOUR
TI - Multimedia Topic Models Considering Burstiness of Local Features
T2 - IEICE TRANSACTIONS on Information
SP - 714
EP - 720
AU - Yang XIE
AU - Koji EGUCHI
PY - 2014
DO - 10.1587/transinf.E97.D.714
JO - IEICE TRANSACTIONS on Information
SN - 1745-1361
VL - E97-D
IS - 4
JA - IEICE TRANSACTIONS on Information
Y1 - April 2014
AB - A number of studies have been conducted on topic modeling for various types of data, including text and image data. We focus particularly on the burstiness of the local features in modeling topics within video data in this paper. Burstiness is a phenomenon that is often discussed for text data. The idea is that if a word is used once in a document, it is more likely to be used again within the document. It is also observed in video data; for example, an object or visual word in video data is more likely to appear repeatedly within the same video data. Based on the idea mentioned above, we propose a new topic model, the Correspondence Dirichlet Compound Multinomial LDA (Corr-DCMLDA), which takes into account the burstiness of the local features in video data. The unknown parameters and latent variables in the model are estimated by conducting a collapsed Gibbs sampling and the hyperparameters are estimated by focusing on the fixed-point iterations. We demonstrate through experimentation on the genre classification of social video data that our model works more effectively than several baselines.
ER -

IEICE TRANSACTIONS on Information