lemna.org hosts genomic data in three categories. Each dataset's own citation file (linked from its download page) states which category it falls under and the exact reference to cite — check there first if you're unsure.
1. Published data
Most data on this site accompany a peer-reviewed publication and are freely available for any use, subject to the citation requirement below. See each dataset's CITATION file for the specific paper.
2. Data from partner institutions, hosted with permission
Some datasets were generated by groups at other institutions, not CSHL, and are mirrored here with the originating group's permission as a convenience to the community. These remain subject to whatever release status and terms the originating group has set for their own data — check the dataset's own citation/documentation for specifics, and direct questions about a specific dataset's status to its own authors, not CSHL.
3. Pre-release data (CSHL Lemnaceae Genome Project)
Some assemblies are made available before publication, in the spirit of the Toronto Statement (Toronto International Data Release Workshop, Nature 461:168–170, 2009): the imperative that community-relevant genomic data be released as early as possible, balanced against the reasonable expectation that the group that generated the data gets to publish the primary analysis without being scooped.
These data are preliminary and may contain errors. By accessing pre-release data, you agree not to publish analyses of these genes or genomic data prior to peer-reviewed publication of a comprehensive genome analysis by the CSHL Lemnaceae Genome Project team and collaborators. This restriction is lifted automatically on that publication (see the dataset's citation file, which is updated once a dataset moves to published status).
What this does and doesn't restrict, while a dataset is pre-release:
- Permitted: using individual genes or short regions for comparative reference (e.g. designing primers, checking whether a gene of interest is present).
- Not permitted: any publication containing genome-scale analysis — annotation, gene families, regulatory elements, repeat content, GC content, whole-genome comparisons, or use of the assembly as a reference for transcriptomic/epigenomic data (RNA-seq, bisulfite-seq, ChIP-seq, etc.) — before the primary publication.
- If you're considering an analysis that might fall under this restriction, or want to discuss collaborative publication, contact the principal investigator before proceeding.
Pre-release assemblies and raw sequence reads may not be redistributed or repackaged during the embargo period.
Licensing and citation
Published CSHL-generated data on this site are released under CC BY 4.0 (attribution required). Partner-institution data hosted here with permission are likewise released under CC BY 4.0, unless that dataset's own documentation specifies more restrictive terms set by the originating group — check there first. We aim for our published data to meet the FAIR principles (Findable, Accessible, Interoperable, Reusable).
To cite data from this site, use the citation given in that dataset's own CITATION file. For pre-release data still under embargo, contact the principal investigator regarding appropriate acknowledgment.
