EMU Interictal Scalp EEG Pipeline¶
What this page tells you
Full pipeline for EMU interictal scalp EEG: export Natus list, copy_emu_db.py to cnt-fs, edit_studylist.py on cnt1 (Admission ID rules), emu_pipeline.py to convert/de-identify/upload to ieeg.org, redcap_emu.py, then delete .mef and Natus files.
Requirements¶
Before proceeding with the pipeline, ensure the following requirements are met:
- Added to IRB
- Access to UPHS workstation "MJ09PSLQ"
- Access to UPHS
isilon-neurologyfolder - Access to PMACS Virtual Desktop Interface (VDI)
- Access to PMACS
cnt-fsWindows HIPAA fileshare (Address:\\172.16.50.149\CNT)- Access to the
/eeg_rawfolder
- Access to the
- Access to PMACS
cnt1Linux HIPAA server- Access to the
/project/eeg_processfolder
- Access to the
- Access to ieeg.org EMU Interictal project folders (e.g.
EMU_Interictal_2024_01to04) - Access to REDCap "EMU Scalp Interictal EEG" project
Workflow¶
Follow these steps to execute the EMU Interictal Scalp EEG Collection, Processing, and Uploading Pipeline:
Note: When going through this processing pipeline, if you ever run a script in cnt1 that attempts to access files in cnt-fs and encounter an issue such as a file or folder not being found, most likely the issue is that your PMACS account authentication between the cnt1 and cnt-fs servers has timed out. Simply run the command kinit in the cnt1 terminal and enter your PMACS password to refresh your connection between cnt1 and cnt-fs.
Step 1: Create EMU Export List¶
- Use Erin's EMU export list pipeline to create an EMU export list.
- Pipeline to generate list of file directories
- Go to EMU Database
- Make view you want in Natus
- In Natus, go to Tools -> Options
- Select columns
- Last Name
- First Name
- Start time
- Duration
- Study name
- CD label
- MRN #
- Birth date
- EEG #
- File path
- Acquired on
- Headbox Type (this is to get the correct channel mapping)
- Create filter for interictal clips
- On left tab, click filters
- Click add
- Name the filter “interictal”
- Select “Study Name”
- Type “interictal”
- Select “Headbox”
- Select All
- Deselect Quantum
- (This makes it exclude Quantum, thus excluding intracranial)
- Alternate list 2024 - for seizures, no filter!
- Export list to Excel
- Tools->Export to Excel
- Filtered Studies
- Save it to
\\isilon-neurology\neurology\projects\EMU interictal clips - Call it
interictal_clip_list(save as an xlsx file)
- Edit list to have full file path
- Run Erin’s Matlab code
prepare_eeg_file_list
- Run Erin’s Matlab code
- Pipeline to generate list of file directories
- Save the export list to the
cnt-fsWindows fileshare. Save it in the folder/eeg_raw/scalp_eeg/1_natus_export/as an.xlsxfile- DO NOT edit the column names in the Excel sheet
Step 2: Login to UPHS Workstation and Copy EMU Database¶
The copy_emu_db.py script copies the EMU Interictal Scalp EEG datasets from the UPHS isilon-neurology folder over to the CNT's PMACS cnt-fs server. The data is in the raw Natus format. It also creates a copy log file.
- Login to UPHS Workstation "MJ09PSLQ" (Erin's UPHS desktop). Login using your UPHS account and password
- In File Explorer, map
cnt-fsWindows fileshare to drive Z: (\\172.16.50.149\CNT). Login using your PMACS account.- Click "Connect using different credentials". For User name, type
PMACS\pennkeyto change the Domain from UPHS to PMACS.
- Click "Connect using different credentials". For User name, type
- Map
\isilon-neurologyfolder to drive Y: (\\isilon-neurology\Neurology). - Open Windows Command Prompt.
- Type "cmd" in the bottom search bar to open
- Run
python copy_emu_db.py Z:/eeg_raw/scalp_eeg/1_natus_export/example_file.xlsx- A CSV file named
interictal_log_XXXX.csvwill be generated in/eeg_raw/scalp_eeg/3_copy_logs/interictal/ - Note: Close the
.xlsxfile before running the Python script. Do not open the file while the process is running.
- A CSV file named
Step 3: Login to VDI and Edit EMU Studylist¶
Pass in the copy log file to the edit_studylist Python script. The script does several things, including figuring out what the Admission ID should be for each recording, what the iEEG.org filename should be, and what the date-shifted start time should be.
-
- Rules for Admission ID:
* "Admission" is defined as a single EMU visit within a 30-day period. Any EMU recordings from a single patient within this time period are given the same Admission ID.
* Recordings that are part of the same "Admission" for a single patient are given the same Admission ID. Patients have different Admission IDs, but if a patient has multiple Admissions, then each Admission is a different Admission ID as well.
- Rules for Admission ID:
-
- The
edit_studylistPython script takes in a previously generated EMU_studylist CSV file. For each record in the current study list being generated, it checks if that record meets the criteria of being within 30 days for the same patient.
* The largest / most recently used Admission ID and REDCap Record ID are needed for the script to know where to continue in generating Admission IDs and REDCap IDs. The script will automatically pull from the REDCap "EMU Scalp Interictal EEG" project the largest existing Admission ID value and REDCap Record ID value.
- The
- Login to the VDI using PMACS credentials. Open the MobaXterm application (search for it)
- SSH into
cnt1using your PMACS account (ssh pennkey@cnt1) - Navigate to
/project/eeg_process/scalp_eeg/scalp_eeg_programs/interictal_pipeline/ - Run
python edit_studylist.py /path/to/current/interictal_log.csv /path/to/previous/studylist.csv- 1. Example path for current interictal log =>
/mnt/cnt-fs/eeg_raw/scalp_eeg/3_copy_logs/interictal/interictal_log_XXXX.csv - 2. Example path for the most recently used EMU Studylist =>
/mnt/cnt-fs/eeg_raw/scalp_eeg/4_studylist/interictal/EMU_studylist_YYYYMMDD_XXXXXX.csv
- 1. Example path for current interictal log =>
- A CSV file named
EMU_studylist_YYYYMMDD_XXXXXX.csvwill be generated in/eeg_raw/scalp_eeg/4_studylist/interictal/- I like to rename the file to append the starting and ending REDCap Record IDs that are a part of this data batch.
- It's important to review the generated EMU Studylist file for any potential errors BEFORE proceeding with the next processing steps. In particular, pay attention that the Admission ID, IEEG Dataset Name, Day, Index, and REDCap Record ID all look correct before proceeding.
- It is helpful to verify that the starting Admission IDs and starting REDCap Record IDs in the generated EMU Studylist file make sense in the context of the larger REDCap database. To check the REDCap database:
- Login to REDCap using your PMACS account. Open the "EMU Scalp Interictal EEG" Project.
- Go to "Record Status Dashboard". Find the most recently used REDCap Record ID (that is, prior to this data batch).
- In Reports, go to the "admission_ids" report. Click through the last few pages and sort by Admission ID to find the newest / largest used Admission ID number (that is, prior to this data batch). Make sure to look through the last few pages and not simply the last page, as sometimes the largest ID number is not in the last page even after sorting.
- It is helpful to verify that the starting Admission IDs and starting REDCap Record IDs in the generated EMU Studylist file make sense in the context of the larger REDCap database. To check the REDCap database:
Step 4: Run EMU Pipeline¶
The emu_pipeline.py pipeline does several things:
-
- It runs
natus2mefto convert the raw Natus recordings into .mef datasets. It uses the headboxes built into the natus converter to figure out the channel mappings. - It looks through the EEG data and finds the gaps where there isn't any recorded data, i.e., it finds the clip times where there is real data. It creates annotations for these clip times and updates the studylist with them.
- It de-identifies the original EEG annotations of PHI and then merges the clip time annotations with the original de-identified annotations.
- Afterward, it uploads the full dataset to iEEG.org into its appropriate iEEG.org project folder. (The iEEG.org EMU Interictal project folders are broken up by year and trimester, for example,
EMU_Interictal_2022_01to04for recordings between January 2022 and April 2022).
- It runs
- Create the ieeg.org project folder (
EMU_Interictal_YEAR_XXtoXX) you are uploading to if it doesn't already exist.- Log in to ieeg.org. Click "Data", then "Create Project". Create "EMU_Interictal_2024_01to04" (example)
- Click on the project name, then click "Open Project". Click "Project Admins". Click Command+F or Ctrl+F, then search for "Erin Conrad". Check the checkbox next to the username. Repeat to grant access to this project for any other collaborators.
- In
cnt1launch a new Linux Screen- Run
screen -S insert_screen_nameto open a new Screen
- Run
- Run
module load pythonto load Python in the new environment - Make sure you're in
/project/eeg_process/scalp_eeg/scalp_eeg_programs/interictal_pipeline/ - Run
python emu_pipeline.py /path/to/emu/studylist.csv- Example path for current EMU studylist file =>
/mnt/cnt-fs/eeg_raw/scalp_eeg/4_studylist/interictal/EMU_studylist_YYYYMMDD_XXXXXX.csv
- Example path for current EMU studylist file =>
- Detach from the Linux Screen. It will take a few minutes to process each record.
- Hold
Ctrl + A + Dto detach from Linux Screen - You can always reattach to a Screen by typing
screen -r insert_screen_name - If you forget the Screen name, you can type
screen -lsto find it
- Hold
- Wait until the batch processing finishes OR the processing halts/freezes
- DO NOT open the EMU Studylist
.csvfile incnt-fswhile this process is being run as the file is actively being modified - If the program is stuck, type
Ctrl + Cto kill the batch processing. You will need to restart the batch processing from where it left off- For the
EMUXXX_DayXX_Xdataset it was in the middle of processing, you will need to delete all of the incomplete files from cnt-fs. Go to the foldercnt-fs/eeg_raw/scalp_eeg/5_mef/EMUXXX/EMUXXX_DayXX_X/and delete its contents - Make a duplicate copy of the current
EMU_studylist_YYYYMMDD_XXXXXX.csvfile incnt-fs. In the duplicate copy, delete all the entries/rows that have already been processed, so that the.csvfile now starts with theEMUXXX_DayXX_Xdataset it was in the middle of processing. Rename this file (I like to append to the filename the starting and ending REDCap Record IDs that are a part of this new batch) - Back in the Linux Screen, run
emu_pipeline.pyagain but using the newly edited EMU Studylist.csvfile
- For the
- DO NOT open the EMU Studylist
- Once the batch has finished processing, logon to ieeg.org and open the
EMU_Interictal_YEAR_XXtoXXproject folder to check if the datasets have uploaded to the ieeg.org platform (user facing)- Typically I just check if the last EEG dataset in the batch is searchable and viewable in the ieeg.org project folder, since the records are processed sequentially
Step 5: Upload to REDCap¶
The redcap_emu.py script uploads the metadata from the EMU Studylist .csv log files to the REDCap "EMU Scalp Interictal EEG" project.
- In
cnt1make sure you're in/project/eeg_process/scalp_eeg/scalp_eeg_programs/interictal_pipeline/ - Run
python redcap_emu.py /path/to/emu/studylist.csv- Example path for current EMU studylist file =>
/mnt/cnt-fs/eeg_raw/scalp_eeg/4_studylist/interictal/EMU_studylist_YYYYMMDD_XXXXXX.csv
- Example path for current EMU studylist file =>
- Logon to REDCap, open the "EMU Scalp Interictal EEG" project, and verify that the records have successfully uploaded.
Step 6: Cleanup cnt-fs¶
I recommend deleting the .mef files from cnt-fs (stored in /eeg_raw/scalp_eeg/5_mef/) to free up disk space (and since these can be regenerated if needed).
- In cnt1 go to
/project/eeg_process/scalp_eeg/scalp_eeg_programs/shared_processing/ - Run
python delete_mefs_from_studylist.py /path/to/emu/studylist.csvto delete the.meffiles that were generated from this batch processing- Example path for current EMU studylist file =>
/mnt/cnt-fs/eeg_raw/scalp_eeg/4_studylist/interictal/EMU_studylist_YYYYMMDD_XXXXXX.csv
- Example path for current EMU studylist file =>
I also recommend deleting the original raw Natus files from cnt-fs to free up disk space (and since these can be re-copied from the UPHS isilon-neurology folder if needed).
- In cnt1 go to
/project/eeg_process/scalp_eeg/scalp_eeg_programs/shared_processing/ - Run
python delete_natus.py /path/to/emu/studylist.csvto delete the Natus files from this batch- Example path for current EMU studylist file =>
/mnt/cnt-fs/eeg_raw/scalp_eeg/4_studylist/interictal/EMU_studylist_YYYYMMDD_XXXXXX.csv
- Example path for current EMU studylist file =>
If you are planning to keep the original raw Natus files, create a yyyymmdd folder in cnt-fs in /eeg_raw/scalp_eeg/2_data_pull/ and move the datasets there.