OJT Extraction
This specialist runs in chat and as a node you can add to a workflow for repeatable, scheduled runs. See the Agent Catalog for how the two relate.
Overview
On-the-job training (OJT) plans hold a lot of structure — the project itself and the tasks that make up its curriculum — but usually as prose or tables inside a document. OJT Extraction reads a training document you attach and pulls that structure into clean, ready-to-use tables.
Ellie produces two things: a project summary that captures the metadata for the training program, and a curriculum breakdown with one row per main task, including its duration and detailed description.
Everything comes straight from the document. Ellie extracts only what's explicitly stated, never inferring, and leaves a field blank when the detail is missing. If you ask without attaching anything, it will politely prompt you to add a document.
What you'll need
- An on-the-job training document — a training plan, module outline, or task-analysis sheet. Word, PDF, or plain text all work.
- Attach the file or add it as context. See Using files in chat and workflows for how attachments work.
How to ask
Attach the document, then ask Ellie to extract it.
Extract the whole plan:
Extract the project details and curriculum tasks from the attached
training document.
Focus on the curriculum:
Pull just the curriculum tasks from this OJT document, with each task's
duration and description.
What you get
Two downloadable spreadsheets:
- Project Summary — the training project's metadata in a single row, such as title, company, approver, location, assignee, dates, and status.
- Project Tasks — one row per main task, with the task title, duration, and detailed description.
Blank fields mean the detail wasn't in the document.
Tips
- Clear section headers and well-formed tables in the source produce the cleanest extraction.
- Nothing is invented — a blank field means the information wasn't present.
- Scanned images need to be converted to text first; text-based documents work best.
- Review both tables against the source before using them downstream.