How Does HIPAA Apply to Privacy and Confidentiality in Healthcare Research?
HIPAA applies to healthcare research when Protected Health Information is used or disclosed by a HIPAA Covered Entity or Business Associate, requiring researchers and institutions to establish an appropriate basis for accessing the information and safeguards for protecting it throughout the research process. Research involving PHI can require individual HIPAA authorization, an approved waiver or alteration of authorization, or another permitted basis for the use or disclosure. Research privacy also extends beyond HIPAA because institutional policies, human-subject protections, data security requirements, and emerging technologies such as artificial intelligence can impose additional controls on how research information is accessed, analyzed, stored, and shared.
Research Consent and HIPAA Authorization Serve Different Functions
Research informed consent and HIPAA authorization address related issues, but one does not automatically substitute for the other. Informed consent addresses participation in research, while HIPAA authorization provides permission for specified uses or disclosures of PHI when authorization is required by the HIPAA Privacy Rule.
A valid HIPAA authorization identifies the information that will be used or disclosed, who is authorized to make the use or disclosure, who may receive the information, and an expiration date or event. Other required information can include the purpose of the use or disclosure.
Organizations can combine research consent and HIPAA authorization into the same document. The combined document still needs to contain the elements required for a valid HIPAA authorization.
Researchers Should Identify the PHI Required for the Study
Researchers should determine which PHI elements are required to conduct the approved research rather than requesting broad access to patient information. The research protocol should establish why particular information is needed and why the research cannot reasonably and effectively be performed without it.
This approach can also be reflected in technical access controls. If a researcher requires diagnoses and laboratory results but does not require patient names or telephone numbers, the research system can restrict access to those unnecessary identifiers.
Restricting access at the system level reduces dependence on individual researchers remembering which information they are permitted to view. The permissions approved during the research review process can become controls within the technology used to retrieve the data.
HIPAA Authorization Can Be Waived or Altered for Research
An IRB or Privacy Board can approve a waiver or alteration of HIPAA authorization when the applicable criteria are satisfied. A waiver can apply to the complete research project or to a particular research activity.
A partial waiver can be relevant when PHI is required for screening or recruitment before obtaining authorization from individuals. A complete waiver can apply in research involving secondary use of existing information when the applicable requirements are satisfied.
Authorization requirements can also be altered. An alteration can modify particular authorization requirements rather than eliminating authorization for the complete research activity.
Impracticability Is Different From Inconvenience
A request for a waiver should address whether the research could practicably be conducted if individual authorization were required. The fact that obtaining authorization would cost more, take longer, or create additional administrative work does not by itself establish that the research is impracticable.
The relevant issue is whether requiring authorization would prevent the research from being practicably conducted. Researchers also need to establish why access to and use of the PHI itself is necessary for the research.
This distinction prevents the waiver process from becoming an administrative shortcut. The analysis concerns whether the research can practicably proceed under the authorization requirement rather than whether authorization is convenient for the research team.
Exempt Research Can Still Raise HIPAA Requirements
A determination that research is exempt from certain human-subject research requirements does not automatically remove HIPAA requirements. The regulatory questions are separate.
An exempt project can still involve PHI held by a HIPAA Covered Entity or Business Associate. For example, researchers might need PHI to identify eligible participants or create a research dataset even when the project receives an exempt research determination.
Activities that are not regulated as human-subject research can also require a separate HIPAA analysis when PHI is involved. The absence of conventional IRB oversight does not establish that PHI can be used or disclosed without considering the HIPAA Privacy Rule.
Research Data Systems Can Enforce Privacy Permissions
Research privacy can be strengthened when the permissions established during regulatory review are incorporated directly into research technology. Instead of giving researchers unrestricted access to an electronic health record, a dedicated research system can provide only the information approved for a particular project.
A researcher performing an initial feasibility assessment might receive aggregate information showing the number of patients meeting specified criteria without receiving identifiable patient records. Additional information can become available after the required research approvals have been obtained.
The same approach can apply to individual data elements. Names, telephone numbers, medical record numbers, or other identifiers can remain inaccessible when they are not required for the approved research.
Controlled Research Tools Can Reduce Unnecessary EHR Access
Providing researchers with access to dedicated research systems can create stronger privacy controls than allowing broad access to complete electronic health records. A full patient record can contain substantially more information than is required for a research question.
Research systems can also provide access to unstructured information when the approved research requires it. Relevant information can exist in pathology reports, radiology reports, discharge summaries, clinical notes, and other free-text records rather than structured database fields.
Controlled search tools can allow researchers to identify relevant information while maintaining authentication requirements and restrictions based on the approved study. This combines research utility with restrictions on unnecessary access.
Research Data Extends Beyond the Medical Record
Human research increasingly uses information from sources outside conventional electronic health records. Genomic information, wearable devices, research repositories, predictive models, and other datasets can provide detailed information about participants.
Wearable devices can generate large volumes of information about movement, physiology, sleep, activity, and other characteristics. These datasets can support machine learning and other forms of analysis while creating requirements for secure storage and controlled access.
The privacy assessment therefore needs to account for the complete research data environment rather than concentrating only on information extracted from the medical record.
De-Identified Data Can Still Create Re-Identification Risks
Information that has been de-identified can still present privacy concerns when it is combined with other datasets. Separate datasets can contain characteristics that allow an individual to be identified through linkage or triangulation even when a single dataset does not directly identify that person.
This issue becomes more relevant as researchers gain access to larger collections of health, genomic, demographic, behavioral, and wearable-device information. Analytical tools can identify relationships across datasets that were difficult to establish when the information was originally collected.
Research institutions can therefore apply safeguards to de-identified information even when the information is not being treated as PHI. De-identification can reduce privacy risk without making every possible future use of the information risk-free.
AI Creates New Research Privacy Questions
AI can analyze large research datasets, develop predictive models, search complex information, and identify relationships that would be difficult to detect manually. These capabilities also create privacy questions concerning where information is processed and what can be inferred from it.
Researchers need to consider whether information entered into an AI system is stored, retained, reused, or accessible to another party. An AI platform suitable for public information is not automatically an appropriate environment for PHI or other sensitive research data.
Controlled systems can allow researchers to develop and validate AI models without submitting sensitive information to open public AI platforms. Access controls, authentication, approved storage, and institutional governance can then operate around the AI research environment.
AI Can Infer Information That Was Never Directly Provided
AI privacy risk is not limited to reproducing information contained in the original dataset. Models can identify patterns and infer characteristics about individuals or groups from combinations of available information.
Possible inferences can concern health status, behavior, demographics, identity, or other sensitive characteristics. A dataset therefore needs to be assessed not only by examining the fields it contains, but also by considering what can be derived from those fields when combined and analyzed.
This creates a different privacy issue from a conventional unauthorized disclosure. The sensitive information might not have existed as an explicit data field before the analysis was performed.
Existing Research Consent May Not Address Future AI Uses
Research consent obtained when information was originally collected may not address later uses involving AI, model training, secondary analysis, or large-scale data sharing. The scope of the original consent therefore matters when researchers consider new uses of existing participant data.
Future AI applications can also be difficult to describe when information is initially collected because the later technology, analytical methods, and research questions may not yet exist. This creates questions about how meaningful consent for broad future use can be when the eventual processing cannot be fully specified.
Research governance needs to consider the permitted use of the information rather than assuming that technical access to an existing dataset establishes permission for every subsequent form of analysis.
Large Research Datasets Can Create Group Privacy Harms
Privacy risks are not limited to identifying an individual participant. Analysis of large datasets can produce findings about communities, demographic groups, or other populations even when no individual person is named.
Research findings can associate a group with a disease, behavior, genetic characteristic, or other sensitive attribute. Those findings can create stigma or disadvantage affecting people who share the characteristic, including people who did not participate in the research.
This expands research privacy analysis beyond the conventional question of whether a particular individual can be identified. The possible effects of data analysis can extend to groups represented within the dataset.
Approved Software and Consumer Software Are Not Necessarily Equivalent
The fact that an organization approves a particular software product does not establish that every account or version of that product is approved for sensitive research information. Institutional versions can operate under different contracts, security configurations, access controls, and data-handling arrangements from consumer accounts.
A researcher using a personal cloud storage account can therefore bypass institutional safeguards even when the same software brand is available through the organization. Research teams need to use the approved service and account rather than relying on familiarity with the product name.
The same distinction applies to AI services, file-sharing platforms, collaboration tools, and other cloud applications. The security and privacy assessment concerns the actual service being used and the conditions under which the information is processed.
Research Information Can Require Safeguards Outside HIPAA
Not all sensitive research information is PHI. Identifiable research data can require institutional safeguards or protections under other laws even when the information is not subject to HIPAA.
Student information provides one example. Research involving identifiable education records can raise FERPA requirements rather than HIPAA requirements, depending on the information and circumstances.
Institutions can also classify research information according to sensitivity and require particular storage, encryption, authentication, or access controls. Researchers therefore need to determine the classification and applicable requirements for the information they handle rather than treating HIPAA status as the only measure of sensitivity.
Human Review Remains Relevant as Research Becomes More Automated
Increasing use of AI does not eliminate the need for human review of research information and analytical results. Researchers still need to determine whether datasets contain identifying information, whether access is permitted, and whether automated findings can be validated against source information.
Automation can reduce manual data extraction, cleaning, and analysis while changing the work performed by research teams. Human involvement can shift toward reviewing data use, validating model outputs, identifying privacy concerns, and determining whether results are supported by the underlying information.
The expansion of AI and large-scale data analysis therefore changes the mechanisms used to conduct research without removing the responsibilities associated with participant privacy, confidentiality, data security, and appropriate access.
