A Theory of Taxonomy
D'Amico, Rabadan, Kleban · physics.soc-ph,cs.SI,physics.data-an,q-bio.QM · 2016-11-04 · 原文
作者:2 位
A taxonomy is a standardized framework to classify and organize items into categories. Hierarchical taxonomies are ubiquitous, ranging from the classification of organisms to the file system on a computer. Characterizing the typical distribution of items within taxonomic categories is an important question with applications in many disciplines. Ecologists have long sought to account for the patterns observed in species-abundance distributions (the number of individuals per species found in some sample), and computer scientists study the distribution of files per directory. Is there a universal statistical distribution describing how many items are typically found in each category in large taxonomies? Here, we analyze a wide array of large, real-world datasets -- including items lost and found on the New York City transit system, library books, and a bacterial microbiome -- and discover such an underlying commonality. A simple, non-parametric branching model that randomly categorizes items and takes as input only the total number of items and the total number of categories successfully reproduces the abundance distributions in these datasets. This result may shed light on patterns i
讲义
讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。
1. 人话版
A taxonomy is a standardized framework to classify and organize items into categories.
Hierarchical taxonomies are ubiquitous, ranging from the classification of organisms to the file system on a computer.
2. 领域脉络
本文类目:physics.soc-ph、cs.SI、physics.data-an、q-bio.QM,属于其所在研究脉络的最新进展。
3. 机制拆解
摘要未展开方法细节——精读时重点看方法/模型部分。
4. 证据与数字
摘要未给出量化结果——留意原文的实验与数据。
5. 反例与边界
Characterizing the typical distribution of items within taxonomic categories is an important question with applications in many disciplines.
Ecologists have long sought to account for the patterns observed in species-abundance distributions (the number of individuals per species found in some sample), and computer scientists study the distribution of files per directory.
Is there a universal statistical distribution describing how many items are typically found in each category in large taxonomies?
A simple, non-parametric branching model that randomly categorizes items and takes as input only the total number of items and the total number of categories successfully reproduces the abundance distributions in these datasets.
6. 跨领域连接与意外收获
横跨 4 个类目(physics.soc-ph、cs.SI、physics.data-an、q-bio.QM),关注其在你兴趣板块间的迁移面。
7. 可复用方法
把本文机制与你手头项目对照,找一个两周内能验证的最小实验。
8. 术语表
精读时把不熟的术语记入此处,作为下次回忆的锚点。