Skip to content

Commit bf01dbc

Browse files
committed
Add DataLad skill for dataset retrieval and computational provenance
Adds skills/datalad covering the Git plus git-annex two-layer model, retrieving content from published datasets, capturing re-executable provenance with datalad run, rerun, and containers-run, publishing to siblings, and the failure modes those involve. Registers the skill in the README catalog and docs/skills.md.
1 parent 48dc1cf commit bf01dbc

6 files changed

Lines changed: 791 additions & 5 deletions

File tree

README.md

Lines changed: 6 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -2,7 +2,7 @@
22

33
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE.md)
44
[![Version](https://img.shields.io/badge/Version-2.64.0-blue.svg)](pyproject.toml)
5-
[![Skills](https://img.shields.io/badge/Skills-163-brightgreen.svg)](#-whats-included)
5+
[![Skills](https://img.shields.io/badge/Skills-164-brightgreen.svg)](#-whats-included)
66
[![Databases](https://img.shields.io/badge/Databases-100%2B-orange.svg)](#-whats-included)
77
[![Agent Skills](https://img.shields.io/badge/Standard-Agent_Skills-blueviolet.svg)](https://agentskills.io/)
88
[![Agent Plugins](https://img.shields.io/badge/Standard-Agent_Plugins-0A7A72.svg)](https://agent-plugins.org/)
@@ -20,7 +20,7 @@
2020
2121
> **Stay up to date:** Follow K-Dense on [X](https://x.com/k_dense_ai), [LinkedIn](https://www.linkedin.com/company/k-dense-inc), [YouTube](https://www.youtube.com/@K-Dense-Inc), and [Reddit](https://www.reddit.com/user/-k-dense-/) for new skills, release announcements, walkthroughs, research workflow demos, and examples you can use with your own AI agent.
2222
23-
A comprehensive collection of **163 ready-to-use scientific and research skills** (covering cancer genomics, individual-level 1000 Genomes queries, hosted regulatory-sequence prediction, live pathogen-variant surveillance, analytical method validation, PK/PD modelling and dose selection, full-text biomedical and regulatory literature retrieval, drug-target binding, bounded biomedical knowledge graph search, molecular dynamics, RNA velocity, microbiome foundation models, geospatial science, time series forecasting, scientific ML resource discovery via Hugging Science, 78+ scientific databases, and more) for any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, created by [K-Dense](https://k-dense.ai). The repository is also a portable [Agent Plugins](https://agent-plugins.org/) package (`plugin.json` + `skills/`), so plugin-capable clients can load the whole collection as one plugin. Works with **Cursor, Claude Code, Codex, Google Antigravity, and more**. Transform your AI agent into a research assistant capable of executing complex multi-step scientific workflows across biology, chemistry, medicine, and beyond.
23+
A comprehensive collection of **164 ready-to-use scientific and research skills** (covering cancer genomics, individual-level 1000 Genomes queries, hosted regulatory-sequence prediction, live pathogen-variant surveillance, analytical method validation, PK/PD modelling and dose selection, full-text biomedical and regulatory literature retrieval, drug-target binding, bounded biomedical knowledge graph search, molecular dynamics, RNA velocity, microbiome foundation models, geospatial science, time series forecasting, scientific ML resource discovery via Hugging Science, 78+ scientific databases, and more) for any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, created by [K-Dense](https://k-dense.ai). The repository is also a portable [Agent Plugins](https://agent-plugins.org/) package (`plugin.json` + `skills/`), so plugin-capable clients can load the whole collection as one plugin. Works with **Cursor, Claude Code, Codex, Google Antigravity, and more**. Transform your AI agent into a research assistant capable of executing complex multi-step scientific workflows across biology, chemistry, medicine, and beyond.
2424

2525
> **Help make AI for science easier to discover:** If Scientific Agent Skills saves you time, teaches your agent a workflow, or helps your lab move faster, please [star this repository](https://github.com/K-Dense-AI/scientific-agent-skills). A star is a public signal that these open, reusable research skills are worth maintaining: it helps scientists, engineers, and open-source contributors find the project, shows which agent-skill standards are gaining real adoption, and gives us a clear reason to keep expanding the collection for the community.
2626
@@ -67,7 +67,7 @@ Recorded walkthroughs of these skills on real research tasks, from the [K-Dense
6767

6868
## 📦 What's Included
6969

70-
This repository provides **163 scientific and research skills** organized into the following categories:
70+
This repository provides **164 scientific and research skills** organized into the following categories:
7171

7272
- **100+ Scientific & Financial Databases** - A unified database-lookup skill provides deterministic, provenance-rich access to 78 public databases (PubChem, ChEMBL, UniProt, COSMIC, ClinicalTrials.gov, FRED, USPTO, and more), plus dedicated skills for DepMap, Imaging Data Commons, PrimeKG, NCATS ARAX, U.S. Treasury Fiscal Data, Hugging Science, OneKGPd, and Genomic Intelligence. Multi-database packages like BioServices (~40 bioinformatics services), BioPython (39 NCBI sub-databases via Entrez), and gget (20+ genomics databases) add further coverage
7373
- **70+ Optimized Python Package Skills** - Explicitly defined, version-aware workflows for RDKit, Scanpy, PyTorch Lightning, scikit-learn, PyTDC, PathML, pydicom, NeuroKit2, PufferLib, QuTiP, GeoPandas, pymatgen, BioPython, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), and others. The agent can still use *any* Python package; these skills provide stronger, safer guidance for the packages listed
@@ -460,7 +460,7 @@ networks, and search GEO for similar patterns.
460460

461461
## 📚 Available Skills
462462

463-
This repository contains **163 scientific and research skills** organized across multiple domains. Each skill provides comprehensive documentation, code examples, and best practices for working with scientific libraries, databases, and tools.
463+
This repository contains **164 scientific and research skills** organized across multiple domains. Each skill provides comprehensive documentation, code examples, and best practices for working with scientific libraries, databases, and tools.
464464

465465
### Skill Categories
466466

@@ -504,8 +504,9 @@ This repository contains **163 scientific and research skills** organized across
504504
- Whole slide imaging: histolab and research-only PathML 3.0.5
505505
- Virtual spatial transcriptomics: noncommercial DeepSpot-M for transcriptome-wide spatial gene expression from 224x224 H&E tiles
506506

507-
#### 🧠 **Neuroscience & Electrophysiology** (3 skills)
507+
#### 🧠 **Neuroscience & Electrophysiology** (4 skills)
508508
- Data standards: BIDS (Brain Imaging Data Structure for neuroscience and biomedical datasets)
509+
- Data distribution and provenance: DataLad (clone and fetch OpenNeuro, DANDI and registry.datalad.org datasets over git-annex, capture re-executable provenance with datalad run/rerun and containers-run, publish to siblings)
509510
- Neural recordings: Neuropixels-Analysis (extracellular spikes, silicon probes, spike sorting)
510511
- Physiological signals: NeuroKit2 0.2.13 for reproducible research workflows—not diagnosis, monitoring decisions, or medical-device validation
511512

docs/skills.md

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -109,6 +109,7 @@
109109

110110
### Neuroscience & Electrophysiology
111111
- **[BIDS](../skills/bids/)** - Brain Imaging Data Structure (BIDS) standard for organizing and describing neuroscience and biomedical research datasets. While originating for MRI, BIDS now covers 11 modalities: imaging (MRI structural/functional/diffusion/perfusion, PET, microscopy), electrophysiology (EEG, MEG, iEEG, EMG), and other data (NIRS, motion capture, behavioral, MR spectroscopy), with active BEPs extending to microelectrode electrophysiology (Neuropixels), stimuli, and more. Covers the BIDS directory hierarchy, file naming conventions with entities (subject, session, task, acquisition, run, etc.), and JSON sidecar metadata. Key features include: dataset creation and validation workflows, querying BIDS datasets with PyBIDS (BIDSLayout), DICOM-to-BIDS conversion using HeuDiConv (ReproIn turnkey, map-into-reproin, and custom heuristic modes), dcm2bids (config-file-based), and BIDScoin (GUI-based), metadata inheritance and sidecar management, events files for task fMRI, participants and scans TSV files, BIDS derivatives conventions for preprocessed data and analysis outputs, BIDS-Apps interface (fMRIPrep, MRIQC, QSIPrep), machine-readable BIDS schema (bids_schema.json) and BEP listing (beps.yml) with update script, and .bidsignore configuration. Includes detailed reference documentation for the complete BIDS specification entity table (35 entities in schema ordering), required and recommended metadata fields for every modality, standard template spaces, and conversion tool workflows with examples. Use cases: organizing neuroscience data for sharing and analysis, validating BIDS compliance before repository submission (OpenNeuro, DANDI), converting DICOM scanner data to BIDS format, creating BIDS-compliant derivatives, querying datasets programmatically, and preparing data for BIDS-Apps processing pipelines
112+
- **[DataLad](../skills/datalad/)** - Retrieve, version, and publish scientific datasets with DataLad 1.6.x over Git and git-annex, and capture computational provenance with `datalad run`, `datalad rerun`, and `datalad containers-run`. Covers the two-layer model that makes a multi-terabyte dataset clone in seconds while holding no data, and the failure it causes when an analysis reads an unfetched pointer instead of a file; finding published datasets through registry.datalad.org, the `///` shortcut to datasets.datalad.org, OpenNeuroDatasets, and dandisets; `clone`, `get -n -r` for subdataset structure, and selective retrieval; git-annex content states inspected with `datalad status --annex`, `git annex whereis`, and `git annex list`; `datalad run` with `--input`, `--output`, placeholders, `--explicit`, `--dry-run`, and the machine-readable run record that `datalad rerun --report`, `--script`, `--since`, and `--onto` read back; `containers-add` and `containers-run` for image-pinned execution via the datalad-container extension; the YODA layout (`datalad create -c yoda`); siblings, `create-sibling-github` and `create-sibling-ria`, RIA `ria+` URLs, and `push --data` semantics with `--publish-depends` to stop publishing history without its content; `drop --what`/`--reckless` for disk management with the deprecated `--nocheck` noted; credential resolution through keyring, `datalad.credential.<name>.<component>`, and `DATALAD_CREDENTIAL_<NAME>_<COMPONENT>`; and `git annex fsck` repair. Needs Python 3.10+ plus a system git-annex 8.20200309 or newer, which is not a pip package. Use cases: fetching OpenNeuro and DANDI data for analysis, pairing with the BIDS skill to run a BIDS-App under recorded provenance, making an analysis re-executable on another machine, and deciding when plain Git is the better choice. Notes the open state of exporting DataLad run records to W3C PROV via datalad-metalad's runprov extractor and BIDS BEP028
112113
- **[NeuroKit2](../skills/neurokit2/)** - Reproducible physiological time-series research with NeuroKit2 0.2.13 on Python 3.10+: method-aware preprocessing, event/interval analysis, multimodal alignment, variability, and complexity across ECG, PPG, EDA, respiration, EMG, EOG, and selected EEG workflows. It is a research and educational toolbox, not for diagnosis, treatment, patient-monitoring alarms, or medical-device validation/certification
113114
- **[Neuropixels-Analysis](../skills/neuropixels-analysis/)** - Comprehensive toolkit for analyzing Neuropixels high-density neural recordings using SpikeInterface, Allen Institute, and International Brain Laboratory (IBL) best practices. Supports the full workflow from raw data to publication-ready curated units. Key features include: data loading from SpikeGLX, Open Ephys, and NWB formats, preprocessing pipelines (highpass filtering, phase shift correction for Neuropixels 1.0, bad channel detection, common average referencing), motion/drift estimation and correction (nonrigid_fast_and_accurate, nonrigid_accurate, and DREDge presets), spike sorting integration (Kilosort4 GPU, SpykingCircus2, Tridesclous2, Mountainsort5 CPU), comprehensive postprocessing (waveform extraction, template computation, spike amplitudes, correlograms, unit locations), quality metrics computation (SNR, ISI violations, presence ratio, amplitude cutoff, drift metrics), automated curation using Allen Institute and IBL criteria with configurable thresholds, model-based curation with pretrained UnitRefine classifiers (noise/neural and SUA/MUA) loaded from Hugging Face via spikeinterface.curation, AI-assisted visual curation for uncertain units using vision-language models, and export to Phy for manual review or NWB for sharing. Supports Neuropixels 1.0 (960 electrodes, 384 channels) and Neuropixels 2.0 (single and 4-shank configurations). Use cases: extracellular electrophysiology analysis, spike sorting from silicon probes, neural population recordings, systems neuroscience research, unit quality assessment, publication-ready neural data processing, and integration of AI-assisted curation for borderline units
114115

0 commit comments

Comments
 (0)