ADP 前沿学习

← 板块一 · 研究前沿

Overcoming data scarcity of Twitter: using tweets as bootstrap with application to autism-related topic content analysis

Beykikhoshk, Arandjelovic, Phung, Venkatesh · cs.SI · 2015-07-10 · 原文

Notwithstanding recent work which has demonstrated the potential of using Twitter messages for content-specific data mining and analysis, the depth of such analysis is inherently limited by the scarcity of data imposed by the 140 character tweet limit. In this paper we describe a novel approach for targeted knowledge exploration which uses tweet content analysis as a preliminary step. This step is used to bootstrap more sophisticated data collection from directly related but much richer content sources. In particular we demonstrate that valuable information can be collected by following URLs included in tweets. We automatically extract content from the corresponding web pages and treating each web page as a document linked to the original tweet show how a temporal topic model based on a hierarchical Dirichlet process can be used to track the evolution of a complex topic structure of a Twitter community. Using autism-related tweets we demonstrate that our method is capable of capturing a much more meaningful picture of information exchange than user-chosen hashtags.

🔮 让 ChatGPT 全网深度追问

讲义

讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。

1. 人话版

Notwithstanding recent work which has demonstrated the potential of using Twitter messages for content-specific data mining and analysis, the depth of such analysis is inherently limited by the scarcity of data imposed by the 140 character tweet limit.

In this paper we describe a novel approach for targeted knowledge exploration which uses tweet content analysis as a preliminary step.

2. 领域脉络

本文类目:cs.SI,属于其所在研究脉络的最新进展。

3. 机制拆解

In particular we demonstrate that valuable information can be collected by following URLs included in tweets.

We automatically extract content from the corresponding web pages and treating each web page as a document linked to the original tweet show how a temporal topic model based on a hierarchical Dirichlet process can be used to track the evolution of a complex topic structure of a Twitter community.

4. 证据与数字

摘要未给出量化结果——留意原文的实验与数据。

5. 反例与边界

This step is used to bootstrap more sophisticated data collection from directly related but much richer content sources.

6. 跨领域连接与意外收获

思考本文机制能否迁移到你正在跟进的问题。

7. 可复用方法

把本文机制与你手头项目对照,找一个两周内能验证的最小实验。

8. 术语表

精读时把不熟的术语记入此处,作为下次回忆的锚点。