vicharanashala-org/Online_Live_Polling_Dataset
Online Live-Polling Dataset Live poll records and attendance from 47 online sessions. The data was collected in an internship programme between May 2026 and July 2026. Sessions were held over Zoom, and the polls were run with Zoom's built-in polling feature, so all tables derive from Zoom's poll and attendance exports. Six tables. sessions.csv and polls.csv describe sessions and polls; the other four record attendance and poll answers at the level of individual participants, who… See the full description on the dataset page: https://huggingface.co/datasets/vicharanashala-org/Online_Live_Polling_Dataset.
Online Live-Polling Dataset
Live poll records and attendance from 47 online sessions. The data was collected in an internship programme between May 2026 and July 2026. Sessions were held over Zoom, and the polls were run with Zoom's built-in polling feature, so all tables derive from Zoom's poll and attendance exports.
Six tables. sessions.csv and polls.csv describe sessions and polls; the other four record attendance and poll answers at the level of individual participants, who appear only as anonymous ids. Tables join on session_id, poll_order and participant_id.
Files
sessions.csv
8 columns, one row per session.
polls.csv
10 columns, one row per poll.
attendance_timeline.csv
3 columns, one row per minute of each session.
attendance_intervals.csv
4 columns, one row per stay. A participant who leaves and rejoins has one row per stay.
poll_responses.csv
4 columns, one row per answer submitted. Which answer was chosen is not included.
participant_sessions.csv
4 columns, one row for every participant who attended a session or answered one of its polls. It is a convenience summary of the two tables above.
Participant ids
Each person is a random id, P0001…P2821, the same in every table and every session. Every account in the Zoom reports is a participant, including the organisers' and speakers' accounts. Ids were assigned in random order and the link between an id and a person was not kept, so an id cannot be traced back to anyone. Names, email addresses, the content of answers, and all dates and clock times are removed; times appear only as offsets within a session.
Reading the columns
Participation is not a column. Divide n_responses by present_at_launch. That works for 603 of the 604 polls; the exception is the one poll nobody answered.
`offset_min_from_start` is measured to the first response, not to the launch. The Zoom export records when each response arrived but never records when a poll was opened, so no launch time exists in the source. The offset overstates when the poll appeared, by however long the fastest respondent took. For the same reason response_offset_s counts from the first response, not from the launch.
`n_distinct_answers` counts distinct answer strings received, not options offered. For a True/False item where both answers were chosen it is 2; where a poll accepted typed answers it can be large, the maximum here being 423. It is 0 for the one poll nobody answered, and 1 where every respondent gave the same answer. The distribution is 534 polls at exactly 2, 45 between 3 and 8, 22 above 8, and 3 below 2. The Zoom export does not record option lists, so the number of options a poll offered is not recoverable.
`attendance_span_min` is not the scheduled length. An attendee who never clicks leave keeps the meeting open, so the span overruns the real session on some days (median 157 minutes, maximum 876). peak_concurrent is the reliable measure of how full a session was.
Participation can exceed 1. present_at_launch is the room size at the poll's first response; attendees who joined during the response window and answered are counted in n_responses but not in the denominator. This happens on one poll (S03 q1: 280 responses against 263 present).
Session `S34` is very small (21 unique attendees, single-digit presence at each of its four polls) and behaves unlike the rest of the corpus.
Every poll has a row, including the one nobody answered, so these are full counts rather than a filtered subset.
The tables are consistent with one another. Counting distinct participants present at each minute in attendance_intervals.csv reproduces attendance_timeline.csv exactly, and from it peak_concurrent and attendance_span_min. Counting rows of poll_responses.csv reproduces n_responses in polls.csv and sessions.csv, and the largest response_offset_s of each poll equals its response_window_s.
`unique_attendees` comes from Zoom's attendee summary report, while attendance_intervals.csv comes from its join-and-leave report. The two reports count slightly differently, so the number of distinct participants in a session's intervals can exceed unique_attendees (by at most 60, and equal in 7 sessions).
Stays. Most participants have several stays in a session (up to 131), since every reconnection starts a new one. 126 stays have zero length.
Participants who never answered. 3,653 rows of participant_sessions.csv have polls_answered 0. Two rows have an empty minutes_present: these participants answered a poll but do not appear in the attendance report. Of the 2,821 participants, 2,419 answered at least one poll.
Poll types
poll_type is assigned by a keyword match on the question text:
Rules are applied in that order, so affective, demographic and icebreaker are checked first. It is a coarse convenience label, and assessment_other is a residual rather than a category.
License
MIT License, see LICENSE.
