How the Paper Collection Is Built

This page explains how the 2,000-paper curated connectomics literature corpus was retrieved, screened, stratified across 12 canonical research domains, structured into 3 nested tiers ($500 \subset 1,000 \subset 2,000$), and annotated with OCAR cards and 3-level pedagogical summaries.


🏛️ Corpus Architecture & Nested Tiers

The collection is structured into three strictly nested materializations:

Tier Corpus Size Primary Target Audience & Role Metadata Depth
500 Key Papers 500 papers Curriculum flagships, course reading lists, seminar deep-dives Full abstract, complete author list, verified 5-part OCAR research cards, 3-level summaries, seminar discussion prompts
1000 Key Papers 1,000 papers Comprehensive scholarly survey, methods reference, subfield tracking Full abstract, complete OCAR cards, citation metrics (in/out degree, k-core), domain classifications, organism tags
2000 Key Papers 2,000 papers Global bibliometric network, citation lineage modeling, AI synthesis Complete directed citation graph ($5,460+$ internal links), complete OCAR cards, full abstract and author/venue metadata, facet views

🔬 Multi-Channel Retrieval & Strict Scope Screening

The candidate pool was compiled using positive nanoscale/synaptic-resolution inclusion gates across Semantic Scholar, OpenAlex, Europe PMC, and PubMed:

  1. Direct Synaptic Connectomics: Dense EM wiring diagrams, synaptic resolution imaging, automated segmentation pipelines (FFN, U-Net, affinity prediction, flood-filling).
  2. First-Class Scientific Axes: Covers tissue preparation, FIB-SEM/SBEM acquisition, synapse detection, proofreading tools (CAVE, CATMAID, FlyWire), graph analysis (motifs, modularity, network topology), structure-function modeling, NeuroAI, cell census, health-translation, and training/outreach.
  3. Positive Nanoscale Boundary: Macroscale non-synaptic methods (such as standard low-resolution fMRI or whole-brain fiber tractography without synaptic validation) are filtered out, preserving a clean nanoscale focus.

📊 Stratified 12-Domain Literature Taxonomy

Candidate papers are classified into 12 mutually exclusive primary domains using a strict decision-order hierarchy:

  1. circuit-structure (15.0% target share / 75 in Top 500 / 300 in Top 2,000)
  2. pipeline (15.0% target share / 75 in Top 500 / 300 in Top 2,000)
  3. physiology (12.0% target share / 60 in Top 500 / 240 in Top 2,000)
  4. behaviour (12.0% target share / 60 in Top 500 / 240 in Top 2,000)
  5. imaging (8.0% target share / 40 in Top 500 / 160 in Top 2,000)
  6. cell-types (8.0% target share / 40 in Top 500 / 160 in Top 2,000)
  7. neuroanatomy (8.0% target share / 40 in Top 500 / 160 in Top 2,000)
  8. synthesis (5.0% target share / 25 in Top 500 / 100 in Top 2,000)
  9. dataset (5.0% target share / 25 in Top 500 / 100 in Top 2,000)
  10. neuroai (5.0% target share / 25 in Top 500 / 100 in Top 2,000)
  11. health (5.0% target share / 25 in Top 500 / 100 in Top 2,000)
  12. training-outreach (2.0% target share / 10 in Top 500 / 40 in Top 2,000)

🃏 What Each Record Carries


🧭 Exploring the Collection