Research data management in Singapore universities is the practical discipline of making sure academic evidence can be understood, protected and responsibly checked after the person who collected it has moved on. A student might finish a compelling thesis, but the result is far less reliable if no one knows which spreadsheet is final, which measurements were changed, who can see participant information or where the approved original records are stored. A data management plan (DMP) helps prevent those failures before they happen.
For students researching university data management plans, NUS research data policy, NTU research repository and reproducible research, the official policies provide a useful foundation. NTU’s Research Data Policy requires project data management planning, controlled storage, preservation of records and responsible deposition or sharing of final research data. NUS Research Data Management says research data generated by staff or students are university property and generally should be kept at least ten years, with more specific registration conditions for defined projects. The rules are not identical.
This guide explains planning, folder structures, metadata, analysis scripts, consent-sensitive information, secure storage, backups, reproducibility and end-of-project handover. It provides non-sensitive fictional examples, not instructions to transfer restricted research files to unapproved services. The student’s university, principal investigator and current agreements determine how real project data must be handled.
The data lifecycle begins before collection
Research data can be planned, collected or generated, cleaned, analysed, described, shared where authorised and retained or disposed of under approved rules. Each stage produces decisions that affect the credibility of the final study.
A student who waits until thesis submission to decide where records belong may discover that the original measurement file is missing or that permissions never covered the intended use. A simple early plan improves both research integrity and practical efficiency.
What counts as research data?
NTU’s policy defines research data broadly, including numerical, descriptive, audio, visual and physical records collected or generated during the project, including models and simulations. The term is not limited to spreadsheet cells or survey answers.
A laboratory image, qualitative interview transcript, programming test log, annotated design or simulated output can all require careful documentation. The appropriate storage and preservation depend on its sensitivity, ownership and research purpose.
What a DMP actually records
A Data Management Plan explains what information will be produced, how it will be stored and described, who can access it, how integrity will be protected and what can be retained or shared at completion. It is a plan for the whole research lifecycle, not merely a form naming a folder.
NTU expressly describes the DMP as the record of intended handling, use and sharing, while NUS Libraries’ data planning workshop describes it as a living blueprint that reduces data loss and decision fatigue.
NTU requires a DMP for research projects
The NTU Research Data Policy applies to staff, researchers, students and others conducting research under its auspices. It requires projects to include a DMP, submitted through the designated university system, and asks the principal investigator to update it when the project changes.
The general obligation should not be mistaken for permission for each student to file data on an arbitrary platform. The responsible PI and current institutional procedure determine submission, review and access. Read the actual scheme and funding requirements before collecting data.
NUS policy is not identical to NTU’s
NUS Research Data Management states that research data generated by NUS staff or students are properties of NUS and should generally be archived for a minimum of ten years. It also describes a special final-research-data register for certain NUS-funded projects.
Do not copy NTU’s data deposition procedure into a NUS thesis or assume that a NUS registry requirement applies to every undergraduate classroom exercise. The university and department’s current rules, funding and data type determine the obligations.
A specified NUS registration rule
NUS asks researchers on NUS-funded studies exceeding S$250,000 per project, with start dates from 1 January 2022 onward, to register final research data supporting published results in the NUS Research Data Register within two weeks of publication of the relevant interim or final result.
This is a defined rule, not a universal instruction that every student upload a thesis dataset publicly within two weeks. A candidate should check whether the project is covered and whether data security or other restrictions apply. Registration and unrestricted access are not the same action.
The principal investigator remains responsible
NTU assigns overall responsibility for effective project data management to the PI, while other research team members have duties under applicable policies and approved procedures. A student should know who is authorised to approve access, storage and sharing.
Do not assume that being the person who made a spreadsheet gives an unrestricted personal right to take it to a new employer or post it online. Research ownership, funding terms and intellectual property agreements matter.
Research ownership is a separate question
NTU’s policy states that the university generally owns data arising under its auspices unless other sponsorship terms, agreements or policies supersede that position. Joint projects require appropriate rights arrangements.
Before contributing to a university-industry project, identify who controls access and which materials can be used for a thesis or portfolio. A student’s legitimate intellectual contribution does not cancel confidentiality or data ownership conditions.
Good file names make evidence traceable
Use descriptive and consistent file names showing project, date or version where the university’s security rules permit. A file called final_final2_revised is likely to cause confusion, especially when several people exchange copies.
For example, a fictional non-sensitive classroom dataset might use studyA_synthetic_v03.csv with a dated notes file explaining changes. Do not place real participant names, identification numbers or protected diagnoses into file names that appear in shared lists.
Separate original records from processed records
Keep original approved source records intact in the designated storage, while making clearly documented copies for cleaning or analysis. Changes to an analysis should not silently overwrite the only surviving record of what was originally observed.
A reproducible workflow can identify which inputs produced a particular result, including the transformations applied. Where access to original identifiable data is restricted, the researcher should follow the authorised controlled environment rather than create personal duplicates.
Keep a simple data dictionary
A data dictionary describes each variable, its unit, permitted values, coding scheme and missing-value treatment. Without it, a future reader may not know whether 1 represents yes, the first group, or a severity category.
For an invented non-sensitive questionnaire, a codebook might explain that response_1 means ‘strongly disagree’ on a specified scale, while a blank value means not answered rather than zero. Accurate documentation prevents analysis errors that ordinary proofreading will not catch.
Metadata explains where data came from
Metadata may record the project, creator, collection date, measurement instrument, method, software and licence or access rules. Its purpose is to help authorised readers interpret a dataset correctly.
NTU’s policy includes formal definitions for metadata and data documentation. A usable repository record requires more than uploading an unexplained file. The correct descriptive information helps others understand whether the data are suitable for verification or legitimate reuse.
Document transformations rather than only outcomes
Data cleaning may involve converting units, handling missing values, merging tables or correcting an observed entry. Those decisions can change the final conclusions, so document the rule and justification.
A student should not remove unusual observations merely because the resulting graph looks less tidy. An outlier might reflect error, a legitimate extreme case or a different group. Reproducibility requires recording how the decision was made.
Version control is about an auditable sequence
For code and non-sensitive project documentation, an institutionally permitted version-control system can record changes and help collaborators know which analysis generated a figure. Version control does not replace secure backup or formal research access control.
A private repository can still be unsuitable for protected personal data if the host, configuration or agreement is not approved. Choose storage and collaboration tools according to university information-security instructions rather than convenience.
Analysis scripts support reproducibility
A saved script or documented set of steps can show how raw approved data were transformed into a table or chart. That is more reliable than a sequence of undocumented manual edits, especially if another analyst must check the result.
A script does not guarantee correctness. Check assumptions, data definitions and results using appropriate tests. When a method depends on random sampling, record the relevant settings where appropriate so a permitted reviewer can reproduce the analysis under comparable conditions.
Reproducibility and replication differ
Reproducibility often concerns obtaining a reported result again from the same data and stated methods. Replication can involve testing the finding in a new study or independent dataset. Both support research integrity but ask different questions.
A student should not claim that rerunning one script proves a result applies universally. The original sample and method may still have limitations. Clear reporting distinguishes an auditable calculation from external evidence that the conclusion generalises.
Secure storage depends on the risk
Research involving identifiable people, commercial information or sensitive technical records may require institutionally controlled encrypted storage and restricted access. The fact that a consumer file-sharing service is convenient does not make it appropriate.
NTU’s research data management resources point researchers to a classification and handling guide. The correct storage plan follows the actual information class and research approvals, not a generic recommendation to upload everything to a personal cloud folder.
Backups must also be authorised
A good backup strategy protects against data loss, device failure and accidental overwriting. But storing duplicate restricted records on a private USB drive or personal account may increase disclosure risk.
Ask the university IT or research support team which approved backup arrangements are available for the project’s classification. A backup is useful only when it preserves availability without undermining confidentiality or institutional control.
Access should follow roles
Not every member of a project needs every record. A data analyst may work with de-identified values while a separate authorised coordinator retains participant contact details. Appropriate separation can reduce exposure if a working folder is shared.
Document who may read, edit or export each category under the project’s approved plan. The student should not grant a new collaborator access merely because it would make communication easier. Confirm permissions with the PI and data owner.
Human-subject consent affects later reuse
Data collected from research participants may be restricted to uses described in their approved consent or other institutional arrangement. A plan to repurpose information for a later thesis or publication can require further review or permission.
Read our research ethics guide for consent and IRB distinctions. A valid analysis technique does not create a right to use identifiable data outside its permitted purpose.
Anonymisation needs more than deleting names
A dataset can remain identifiable through rare combinations of details, linked identifiers, free-text comments or unusual events. A researcher should assess the realistic possibility of re-identification rather than assume that removing a name makes all records anonymous.
Pseudonymised data with a separate re-identification key still require protection. Follow the approved institutional classification and access rules, especially when results might be shared beyond the original team.
Open sharing is not the same as sharing everything
Universities encourage data accessibility and verification, but legal, ethical and commercial conditions can restrict public release. NTU’s policy explicitly recognises these constraints and requires sensitive data to remain in suitable secure arrangements.
A defensible project may share a codebook, analysis script or permitted synthetic dataset while withholding identifiable records. The correct decision is set by the data owner, funder, IRB and institutional rules, not by an individual’s preference for maximum openness.
NTU’s final-data deposition requirement
NTU’s data policy requires final research data to be deposited in its research data repository or an appropriate external repository, subject to responsible restrictions, with deposit timed to publication or project completion under the policy’s conditions.
It also requires registration of externally hosted datasets and documentation. An academic should not assume this means publishing identifiable interviews openly. Repository submission, metadata availability and file-access restrictions can differ.
Retention is not ‘keep it forever’
NTU states a general ten-year retention period following publication or project completion, whichever is later, with extensions where proceedings or other obligations require it. NUS also generally requires retaining research data for at least ten years.
Those are institutional baselines, not permission to disregard a particular sponsor’s longer period or a legally required secure deletion policy. A student should understand who remains responsible after they graduate and which records must be preserved in controlled storage.
A README is a low-cost research improvement
A readable README can describe what each folder contains, the meaning of files, which version is authoritative, how analyses are run and whom an authorised researcher should contact. It supports a new student joining the project or a supervisor checking a figure.
For a fictional synthetic dataset, the README might state the file format, number of rows, expected variables and precise command for generating a chart. It should not publish secrets, participant identifiers or protected network information.
Use consistent units and definitions
Research can fail quietly when one table records seconds and another milliseconds, or when two team members use a variable label differently. A data dictionary should explain units, category meanings and valid ranges.
An engineering student could record temperature in degrees Celsius with the instrument and sampling interval; a business student might specify whether an amount is gross revenue or cost. Explicit definitions prevent apparently precise calculations from answering the wrong question.
Missing values are not automatically zero
A blank survey response could mean not asked, participant declined, data lost or genuinely zero. Treating all missing information as zero can distort a summary or model.
Document the reason where known and use a defensible handling method. The analysis may need to retain missingness or compare results under alternative approaches. Never silently replace gaps with convenient numbers that make a thesis conclusion more attractive.
Why a data cleaning log matters
Data cleaning is not merely cosmetic. Correcting duplicates, inconsistent dates or measurement errors can change sample size and statistical conclusions. Record the method and who authorised changes.
A short log might state that two records were duplicates under a defined identifier, or a sensor result failed a specified quality check. A reviewer should be able to distinguish justified cleaning from deleting inconvenient results. Transparency supports research integrity.
Document software and computation environment
A numerical analysis can depend on software versions, libraries and random-number settings. When permitted, record those dependencies and the commands needed to recreate outputs. A small environment note can save substantial time later.
Do not claim that matching computer output proves the entire scientific explanation. It shows the calculation can be reproduced under documented conditions. The soundness of data collection, model assumptions and interpretation still need independent evaluation.
Synthetic data can support safe learning
A fictional or randomly generated dataset can help students learn cleaning, plotting and model checks without risking real participant privacy. Mark every synthetic file clearly and avoid mixing it with approved research observations.
An analysis conducted only on invented data cannot be presented as empirical evidence about real people or institutions. Use it to demonstrate procedure, test code or clarify a research plan. Actual findings must come from appropriately collected and authorised evidence.
A fictional computing project
Imagine a computing postgraduate student testing a classifier using a synthetic dataset. They save an unmodified input file, a script describing preprocessing, a test report and a README that identifies which run produced each result.
A later reader can inspect assumptions and reproduce the calculation. But the student must still discuss whether the synthetic data reflect any relevant real-world phenomenon. Reproducible code and externally valid conclusions are different standards.
A fictional engineering laboratory project
Imagine an engineering researcher measuring a non-hazardous classroom quantity under authorised supervision. A usable record would identify the instrument, calibration information, timestamp, unit, observed readings and any deviation from the method.
A figure without those details may look polished while concealing measurement uncertainty. Real hazardous or regulated work must follow approved laboratory protocols and safety procedures. The educational example is about documentation, not technical permission to operate equipment.
A fictional interview study
Imagine a researcher studying adult learning experiences through approved interviews. Identifiable recordings are stored securely under the protocol, while an authorised de-identified working document is used for thematic analysis.
The analysis should record how themes were defined and which quotations can be used. It must not reveal participants through unique contextual details. The IRB-approved consent and data agreement decide what can be disclosed, even when a result appears especially interesting.
A fictional economic dataset
An economics student combines two public statistical series. Their DMP identifies source URLs, publication dates, units, revisions and the procedure for joining observations across time. The final report explains which years are comparable.
A new edition of a statistical table could alter values, so the student should record the source version used. Data provenance makes a conclusion easier to verify and prevents a later reader from assuming figures taken from different periods are identical.
Public repositories need appropriate licences
When data can be shared, a repository record may indicate a licence controlling reuse and attribution. Researchers should choose a licence that matches institutional ownership, participant permissions, funder rules and any third-party material.
A student cannot attach an unrestricted licence to data owned by an external company without authority. The repository administrator or university library can explain permitted options. Publication and download access should remain within the project’s approved constraints.
Repository metadata can remain visible when files are restricted
An institutional repository can record a project’s existence and descriptive information even when sensitive data are under controlled access or embargo. This can support discoverability without exposing protected records.
Check which metadata fields themselves might reveal sensitive information. A vague description may be inadequate for reuse, but an excessively detailed public record can also create privacy risks. The institution’s data team can help balance those concerns.
NTU offers a dedicated data repository
The NTU data management page describes DR-NTU (Data), a repository that curates and preserves research outputs, supports funder requirements and enables responsible sharing. The university also provides training and a data classification and handling guide.
Using the institution’s authorised workflow is not the same as simply emailing all files to the library. The project PI and data owner should determine what final data and documentation may be deposited, and whether access limitations or external repository links are needed.
NUS also provides specialised research support
NUS Libraries offers guidance on data management planning and scholarly communication, while the NUS research office publishes institutional retention and registration information. A student can use authorised university services to clarify file organisation or permitted sharing.
This is preferable to assuming that a generic online storage recommendation applies to confidential interviews, sensitive research or sponsor-owned datasets. Seek practical advice early enough to improve the plan before large volumes of evidence have accumulated.
An academic departure is a real data-management event
When a student graduates, changes supervisor or leaves a laboratory, others may need to access project records for verification or continuing research. Transfer responsibilities and permissions should be resolved through the PI and university rather than by copying everything onto a personal device.
Prepare a permitted handover record showing where final files reside, which methods created outputs and which obligations remain. Ensure unauthorised personal copies are handled according to the approved plan. Ownership and confidentiality can continue long after enrolment ends.
Responding to a suspected data error
If a researcher discovers that an important file has been altered or a measurement may be wrong, the right response is to preserve appropriate evidence and raise the concern with authorised supervisors or research integrity channels.
Do not quietly replace a published result or delete records to hide a discrepancy. Investigate with a documented process, revise conclusions where warranted and follow institutional reporting requirements. Transparent correction protects the credibility of a project.
Data leakage needs an incident route
An accidental disclosure or loss of participant information can create obligations under the university’s policy and applicable law. The student should know the proper reporting contacts and follow their instructions promptly.
Do not use public social media to seek help while revealing the very records at risk. Safeguard further access if authorised, preserve relevant facts and contact the responsible research or information-security office. Timely, accurate reporting is preferable to concealment.
Responsible AI use and research data
AI tools may help with permitted coding or writing tasks, but uploading unpublished research, confidential documents or participant data to an external service can violate the applicable data agreement. Anonymisation and permitted processing must be checked before any transfer.
Students should use institutionally approved tools, follow relevant AI-use and integrity policies, and verify generated analyses or citations. An AI-produced answer is not itself research evidence. The research team remains responsible for what is submitted and published.
A small team needs explicit workflow ownership
In a group project, decide who receives source data, who cleans or analyses them, who checks results, and who maintains the approved repository. These responsibilities can be recorded without disclosing private credentials.
An unexplained shared folder makes it difficult to know who changed a file or whether an output came from the current model. Assigning roles and version conventions improves collaboration and reduces avoidable conflict about contributions.
Make the DMP useful during the semester
A data management plan should be revisited when new variables, collaborators, tools or sources are introduced. NTU expressly requires updates to reflect substantive project changes. This turns the plan into an operational document rather than paperwork filed and forgotten.
A monthly review could check access, backups, changes to data collection, approved storage and pending sharing decisions. It should use the project’s actual risk level and procedures. Some projects need more frequent monitoring than a simple non-sensitive classroom dataset.
A practical handover checklist
At completion, identify which data are final, document the cleaning and analysis workflow, confirm the correct storage location, check repository or registration obligations, and retain the approvals and agreements the institution requires.
Then verify that the appropriate supervisor or PI can access the authoritative record without relying on the graduating student’s private password. A good handover keeps legitimate research usable without extending unauthorised access.
Frequently asked questions
Does a DMP mean all files must be public? No; permissions and sensitive-data restrictions matter. Can personal cloud storage be used for any thesis? Not automatically. Does deleting participant names always anonymise data? No; re-identification can remain possible.
Do NUS and NTU use identical deposit rules? No. What is the purpose of retaining research data? Verification, integrity and lawful future use. Is a reproducible script proof of scientific validity? No; data quality and assumptions still matter.
Official sources and connected eduKate reading
Start with NTU Research Data Policy, NTU Data Management Support, NUS Research Data Management and NUS DMP training. The institution and project agreement govern real work.
Related eduKate guides include Research Ethics Review, Choosing a PhD Supervisor and Academic Research Writing.
Final thought: evidence should outlive the researcher
Strong research data management makes work easier to understand, review and preserve while protecting people and partners whose information was entrusted to the project. The aim is not to create paperwork for its own sake but to make important conclusions traceable to reliable evidence.
Start with a realistic DMP, use approved storage, document transformations and finish with a proper handover. When research records are trustworthy, the eventual thesis or publication has a stronger foundation. Reviewed 8 October 2026; current NUS, NTU, legal and funding rules remain authoritative.
