1  Syllabus

12-778 | Sec. A | Fall 2026

TR 2:00PM - 3:20 PM ET Location: PH A7F


Instructors: Mario Bergés, PH 123L, Phone: x8-4572 Office Hours: Tuesdays 1:00 PM - 2:00 PM ET in PH 123G and by appointment.

Teaching Assistant: Avi Dube (avid@andrew.cmu.edu) Office Hours: Tuesdays 1:00 PM - 2:00 PM ET in 118G and by appointment.

Textbooks (optional):

  1. Though no official textbooks are required, the instructor will draw from multiple sources and post the relevant references on Canvas.

Prerequisites: Probability & Statistics, Linear Algebra

1.1 Official Course Description

AI and machine learning are rapidly becoming standard tools for civil and environmental engineering, but their performance depends on how data are measured, stored, and connected to the physical systems they represent. This course provides the end-to-end foundations needed to develop reliable data-driven solutions for engineering problems (from sensing and digitization to data management and modeling) serving as a bridge into graduate machine learning. The course begins with measurement physics and instrumentation, studying how physical processes are transduced into electrical signals, conditioned and filtered, and converted to digital information. Topics include common sensor modalities, interface circuits, DC/AC circuit analysis, amplification, filtering, frequency-domain analysis, and analog-to-digital conversion. The course then develops the “data layer” required for scalable deployments, including uncertainty and measurement error, compression and representation, metadata, and database design. Finally, students learn modeling and inference concepts that prepare them for ML across learning paradigms, including supervised learning foundations (regression, regularization, validation, and generalization) and unsupervised methods for structure discovery and monitoring (dimensionality reduction, clustering, and anomaly detection). A final project integrates these components to design a sensing-to-modeling pipeline for a civil or environmental engineering application, emphasizing domain knowledge and physical consistency throughout the designed solution.

2 Grading

Component Weight
Assignments 20%
Mid-term Exam 10%
Progress Update Report 20%
Final Exam 20%
Final Project Demo/Poster 20%
Knowledge Contributions 10%

2.1 Assignments

A total of four assignments are scheduled. The topics covered in each assignment will closely follow the concepts covered in class. Beyond exercising these concepts, the assignments are designed to progressively build the practical skills you will need for the final project: setting up and using sensing hardware, programming microcontrollers, designing and populating databases, and training models.

All assignments are to be solved individually. Discussions and conversations with other students regarding the problem sets are encouraged. However, the final solutions along with the reasoning behind them need to come from you and be clearly explained in the submitted documents.

Each assignment will be worth 5% and is due at the beginning of class on the date that is indicated in each assignment. Assignments that are submitted before this deadline can receive 100% of the available credit. There is a 3 day grace period for late submissions, with the following available credit in each of those three days: 1 Day Late: 90%, 2 Days late: 70%, 3 Days late: 40%. After this time, assignments will not be graded. Of course, if you anticipate not being able to meet this schedule due to a major problem, please talk to the instructor as soon as possible.

2.2 Written Knowledge Contributions

A small portion (10%) of the grade for the course will be based on the individual contributions each of you make to the body of knowledge we are collectively capturing on a website/book for the course. The main way you will gain points here is by suggesting corrections for errors, omissions or other possible improvements to the content. I’m running an experiment to test whether pre-trained LLMs from frontier labs can generate the content for the lecture notes. This is the third time I run this experiment, but it is still far from perfect. I expect there will be multiple opportunities to spot errors. The other way to contribute and earn these points is by submitting your own content: terms, definitions, methods, summaries, worked-out exercises, diagrams, and other contributions are fair game for this purpose. The website is hosted on a private Gitlab instance, and will require that you log in with an account in order to edit it. This will give me the opportunity to evaluate your contributions by simply following the commit log of the repository.

2.3 Project Update

A fifth (20%) of the grade for the course will be based on a progress update that will take place during the second half of the course. This update will consist of: (1) a written 2-page report, authored by all team members; and (2) a 5 minute individual meeting with the instructor to discuss the project, its goals and the plan forward. The written progress report will confer 10% of the final grade, while the individual meeting discussion will be used to provide the remaining 10%.

2.4 Final Project Demo/Poster

The final project culminates in an in-person demo/poster session held on the last day of classes (Friday, December 4). It is worth a fifth of the total grade (20%) and, in some ways, is the most important assignment. There is no written report: the poster and the accompanying demonstration are the deliverables. We recommend a 36”×48” poster (the standard size for academic conference posters). Use your own words when preparing them and avoid plagiarism of any kind. On the poster, you will describe, in simple terms, the motivation and specific objectives of your project, the design choices made for the hardware and/or software prototype that you put together, the experiments you performed to validate whether or not your solution satisfies the objectives, and a discussion about the results and limitations. You should make all code and datasets available as part of your submission, along with a digital copy of the poster, by the end of the day on December 4.

The rubric that will be used for grading the poster and demo is as follows:

  • Formatting and organization of the poster (15%)
  • Grammar, clarity and accuracy of ideas (15%)
  • Display of mastery of concepts covered in class (35%)
  • Creativity expressed in the solution (10%)
  • Depth of discussion related to the pros/cons of the implemented solution (15%)
  • Overall assessment of the project’s idea and execution (10%)

3 Course Policies

Though there are definitely reasons to like anarchism, I still prefer democracy, so here are the rules of the game as they are now (and subject to change if enough of you request that I do).

3.1 Collaboration

Collaboration is expected within the limits of discussing concepts and problems. However, each student must produce his/her own solution to the problems. Copying from another student’s assignment is clearly plagiarism. Using information directly from websites, books, papers and other literary sources without appropriate attribution is also plagiarism. Assignments submitted for this class will be reviewed by the instructor and TA and may be scanned through web-based academic integrity software. Occurrences of cheating or plagiarism will be handled according to the university policy on Academic Integrity, https://www.cmu.edu/policies/documents/Academic%20Integrity.htm. Students are expected to have read this policy and conform to the highest standards of academic integrity. For incidents of academic misconduct, the University Academic Disciplinary Actions Policy, found at https://www.cmu.edu/student-affairs/theword/acad_standards/creative/disciplinary.html, will be followed.

3.2 Use of AI

Collaboration is no longer limited to the human-human pair, so I need to address collaboration with AI systems separately.

You are welcome to use generative AI systems in this course. Just like having a smart and dedicated classmate to study and work together with, AI systems can enhance your learning, cut down the time you need to take to deliver assignments, and become more productive overall. However, your ethical responsibilities as a student remain the same. Crucially, you want to ensure that you leave this class with independent command of the concepts we are covering in class. If you want to be attractive in the job market, you clearly need to be able to deploy this newly acquired knowledge more effectively than the AI system alone (or your smart and dedicated classmate, for that matter).

You must follow CMU’s academic integrity policy. Note that this policy applies to all uncited or improperly cited use of content, whether that work is created by human beings alone or in collaboration with a generative AI. If you use a generative AI tool to develop content for an assignment, you are required to cite the tool’s contribution to your work. In practice, cutting and pasting content from any source without citation is plagiarism. Likewise, paraphrasing content from a generative AI without citation is plagiarism. Similarly, using any generative AI tool without appropriate acknowledgement will be treated as plagiarism.

3.3 Class Participation

Students are expected to be in class on time and participate in class discussions. If you cannot make class, please inform your instructors and group members ahead of time. In class, students are expected to be courteous and respectful of the views and needs of other students and instructors.

3.4 Students with disabilities

Students requesting classroom accommodation must first register with the Dean of Students Office. The Dean of Students Office will provide documentation to the student who must then provide this documentation to the Instructor when requesting accommodation.

3.5 Posting of course materials

All the material used in the course (syllabus, readings, problem sets, reports) is intended for use in the class only. No unauthorized posting, publication or redistribution is expected. Uploading course materials to Course Hero or other web sites is not an authorized use of the course material.

4 Course Outline

Here is a more detailed overview of how the lectures will unfold over the three periods. Most of this is subject to change. The list of “learning objectives” should be understood as responding to the phrase By the end of the lecture, students should be able to

4.1 First third: Sensors

Learning Objectives:

  • Recognize the relative importance of the assessed components of the course (final project, assignments and written contributions)
  • Get to know other students in the course and their motivation for joining it
  • Understand the goals of the course and its structure

References:

No references for this week.

Learning Objectives:

  • Describe the measurement chain that links a physical quantity of interest (the measurand) to a recorded value: transduction, conditioning, digitization, and storage
  • Classify common transduction principles (resistive, capacitive, inductive, piezoelectric, thermoelectric, optical) and give an example sensor for each
  • Interpret the static characteristics reported in a sensor datasheet: sensitivity, range, resolution, linearity, hysteresis, and drift
  • Select an appropriate sensing modality for common measurands in civil and environmental engineering (strain, displacement, acceleration, temperature, humidity, flow, water quality)
  • Explain how the physics of a transducer constrains the quality and interpretation of the data it produces

References:

  • Fraden, Jacob. Handbook of Modern Sensors: Physics, Designs, and Applications, 5th ed. Springer, 2016. (Chs. 1–4)
  • Morris, Alan S., and Reza Langari. Measurement and Instrumentation: Theory and Application, 2nd ed. Academic Press, 2015. (Chs. 1–2)
  • Figliola, Richard S., and Donald E. Beasley. Theory and Design for Mechanical Measurements, 7th ed. Wiley, 2019. (Chs. 1–2)

Learning Objectives:

  • Apply Ohm’s law and Kirchhoff’s voltage and current laws to analyze simple resistive networks
  • Analyze voltage dividers and Wheatstone bridge circuits, and compute the output voltage of a bridge with one or more active sensing elements (e.g., strain gauges)
  • Use complex impedance and phasors to analyze simple AC circuits containing resistors, capacitors, and inductors
  • Derive the time- and frequency-domain behavior of first-order RC circuits
  • Reduce a circuit to its Thévenin equivalent and explain why source and input impedances matter when connecting sensors to instruments

References:

  • Alexander, Charles K., and Matthew N. O. Sadiku. Fundamentals of Electric Circuits, 7th ed. McGraw-Hill, 2020. (Chs. 2–4, 9)
  • Horowitz, Paul, and Winfield Hill. The Art of Electronics, 3rd ed. Cambridge University Press, 2015. (Ch. 1)

Learning Objectives:

  • Identify the components of a data acquisition (DAQ) pipeline and the role each plays: sensor, signal conditioning, multiplexer, sample-and-hold, analog-to-digital converter
  • Explain how an analog-to-digital converter maps a continuous voltage to a discrete code, and compute the resolution implied by a bit depth and full-scale range
  • Quantify quantization error and its contribution to the signal-to-noise ratio of a measurement
  • Distinguish single-ended from differential measurements and identify when each is appropriate
  • Specify DAQ requirements (bit depth, input range, sampling rate, channel count) for a given sensing application
  • Describe how networked sensing architectures (IoT devices, wireless sensor networks) extend the single-instrument DAQ pipeline, and the constraints (power, bandwidth, synchronization) they introduce

References:

  • Taylor, H. Rosemary. Data Acquisition for Sensor Systems. Springer, 1997.
  • Fraden, Jacob. Handbook of Modern Sensors, 5th ed. Springer, 2016. (Ch. 5)

Learning Objectives:

  • Explain the purposes of signal conditioning: amplification, level shifting, isolation, and filtering
  • Analyze basic operational amplifier configurations (inverting, non-inverting, differential/instrumentation amplifiers) under the ideal op-amp assumptions
  • State the Nyquist-Shannon sampling theorem and predict the aliased frequency of an undersampled tone
  • Explain why anti-aliasing filters must be analog and placed before the ADC, and design a first-order low-pass filter for that purpose
  • Choose a sampling rate for a given application, balancing bandwidth, storage, and aliasing concerns

References:

  • Horowitz, Paul, and Winfield Hill. The Art of Electronics, 3rd ed. Cambridge University Press, 2015. (Ch. 4)
  • Lyons, Richard G. Understanding Digital Signal Processing, 3rd ed. Prentice Hall, 2010. (Chs. 1–2)

Learning Objectives:

  • Explain, conceptually, how the Fourier series and Fourier transform decompose a signal into sinusoidal components
  • Compute and interpret the discrete Fourier transform (DFT) of a sampled signal, relating bin indices to physical frequencies
  • Recognize spectral leakage and explain how windowing mitigates it
  • Estimate and interpret the power spectral density of a stationary signal
  • Explain how a linear time-invariant system is characterized by its frequency response, and use transfer functions to reason about the dynamics of sensors and filters
  • Apply the FFT to real measurement data (e.g., identify the dominant modal frequency of a vibrating structure from accelerometer data)

References:

  • Oppenheim, Alan V., and Alan S. Willsky. Signals and Systems, 2nd ed. Prentice Hall, 1996. (Chs. 3–5)
  • Lyons, Richard G. Understanding Digital Signal Processing, 3rd ed. Prentice Hall, 2010. (Chs. 3–4)

Learning Objectives:

  • Distinguish systematic from random errors, and accuracy from precision, in a measurement system
  • Identify common noise sources in instrumentation (thermal noise, interference, quantization) and strategies to mitigate them
  • Explain the role of calibration in reducing systematic error
  • Quantify measurement uncertainty from repeated observations (Type A) and from specifications or prior knowledge (Type B)
  • Propagate uncertainty through a functional relationship using first-order (Taylor series) methods, and explain why propagation for linear Gaussian models is exact while non-linear models require linearization or Monte Carlo approaches
  • Combine independent uncertainty components and report a measurement result with a defensible uncertainty statement

References:

  • Taylor, John R. An Introduction to Error Analysis, 2nd ed. University Science Books, 1997.
  • Holman, Jack P. Experimental Methods for Engineers, 8th ed. McGraw-Hill, 2011. (Ch. 3)
  • JCGM 100:2008. Evaluation of Measurement Data — Guide to the Expression of Uncertainty in Measurement (GUM). https://www.bipm.org/en/committees/jc/jcgm/publications

4.2 Second third: Data

Learning Objectives:

  • Represent time-series data unambiguously, handling timestamps, time zones, and irregular sampling
  • Decompose a time series into trend, seasonal, and residual components to reveal patterns in sensor data
  • Resample, interpolate, and align time series recorded at different rates or with gaps, using regression to estimate unobserved values
  • Detect and handle missing values and outliers in sensor data (e.g., using Chauvenet’s criterion), and articulate the assumptions each strategy makes
  • Explain what virtual sensors are and how they derive measurements from other sensors or models
  • Compute rolling statistics and apply simple smoothing filters to noisy measurements
  • Compare storage formats for time-series data (CSV, Parquet, purpose-built time-series databases) in terms of size, speed, and interoperability
  • Use pandas to carry out all of the above on real sensor data

References:

  • McKinney, Wes. Python for Data Analysis, 3rd ed. O’Reilly, 2022. (Ch. 11) Open edition: https://wesmckinney.com/book/
  • Hyndman, Rob J., and George Athanasopoulos. Forecasting: Principles and Practice, 3rd ed. OTexts, 2021. (Ch. 3) Free at https://otexts.com/fpp3/

Learning Objectives:

  • Define sets and apply the basic set operations (union, intersection, difference, complement, Cartesian product) and their properties, and explain how set theory underpins the relational data model
  • Identify entities, attributes, relationships, and cardinalities — including weak entity sets, roles, and subclasses — in a description of an engineering system
  • Construct an entity-relationship diagram (ERD) for a sensing deployment (e.g., buildings, floors, sensors, measurements)
  • Relate the components of the relational model (relations, attributes, tuples, types) to their table counterparts (tables, columns, rows, domains), and distinguish a schema from an instance
  • Apply and compose the core relational algebra operators (select, project, join, union, difference, rename) to express queries
  • Translate an ERD into a relational schema with appropriate primary and foreign keys

References:

  • Garcia-Molina, Hector, Jeffrey D. Ullman, and Jennifer Widom. Database Systems: The Complete Book, 2nd ed. Prentice Hall, 2009. (Chs. 2, 4)

Learning Objectives:

  • Distinguish the Data Definition Language (DDL) and Data Manipulation Language (DML) families of SQL statements, and create tables, insert, update, and delete records with them
  • Write SELECT queries with filtering, ordering, aggregation, and GROUP BY
  • Combine data from multiple tables using joins, and explain the difference between inner and outer joins
  • Use subqueries and nested SELECT statements to answer multi-step questions
  • Translate between relational algebra expressions and their equivalent SQL queries
  • Set up and issue queries against a local DBMS (e.g., SQLite or MySQL), including from Python (sqlite3, SQLAlchemy) with results loaded into pandas

References:

  • Garcia-Molina, Hector, Jeffrey D. Ullman, and Jennifer Widom. Database Systems: The Complete Book, 2nd ed. Prentice Hall, 2009. (Ch. 6)

Learning Objectives:

  • Identify insertion, update, and deletion anomalies in a poorly designed schema
  • Derive functional dependencies from the properties of the data, and compute closures and keys of relations
  • Normalize a schema through the normal forms, decomposing a relation up to Boyce-Codd Normal Form (BCNF) when possible
  • Explain what transactions and the ACID properties guarantee, and why they matter for concurrent data collection
  • Describe how indexes speed up queries and what they cost, and weigh normalization against denormalization for high-volume time-series workloads

References:

  • Garcia-Molina, Hector, Jeffrey D. Ullman, and Jennifer Widom. Database Systems: The Complete Book, 2nd ed. Prentice Hall, 2009. (Ch. 3)

Learning Objectives:

  • Explain how integers, floating-point numbers, and ASCII text are represented in binary, and predict the size of a simple file from its content
  • Define information entropy and use it to bound the compressibility of a data source
  • Apply Huffman coding to an arbitrary string, and describe other common lossless techniques (run-length, dictionary-based) and where each performs well
  • Contrast lossless and lossy compression, and explain how frequency-domain (transform) representations enable lossy schemes such as JPEG
  • Apply delta and delta-of-delta encoding ideas to time-series data and estimate the achievable savings
  • Debate whether large language models are, in a meaningful sense, compression algorithms

References:

  • Sayood, Khalid. Introduction to Data Compression, 5th ed. Morgan Kaufmann, 2017. (Chs. 1–3)

Learning Objectives:

  • Recognize the workloads for which the relational model becomes awkward, and the alternatives that emerged in response
  • Describe the key-value, document, columnar, object, and graph data models, with an example system for each
  • Explain, at a high level, the trade-offs distributed data stores make between consistency, availability, and partition tolerance
  • Describe what purpose-built time-series databases optimize for
  • Select an appropriate storage model for a given sensing deployment and justify the choice

References:

  • Silberschatz, Abraham, Henry F. Korth, and S. Sudarshan. Database System Concepts, 7th ed. McGraw-Hill, 2019. (Ch. 8)
  • Robinson, Ian, Jim Webber, and Emil Eifrem. Graph Databases, 2nd ed. O’Reilly, 2015. (Chs. 1–2)

Learning Objectives:

  • Explain why metadata, not sensor data itself, is often the bottleneck for scaling data-driven applications across buildings and infrastructure systems
  • Represent facts as subject-predicate-object triples using the Resource Description Framework (RDF)
  • Explain the role of ontologies and controlled vocabularies in making metadata machine-actionable
  • Write simple SPARQL queries against an RDF graph
  • Use a domain schema such as Brick to describe the sensors, equipment, and relationships in a real building

References:

  • Balaji, Bharathan, et al. “Brick: Towards a Unified Metadata Schema for Buildings.” Proceedings of ACM BuildSys 2016. https://doi.org/10.1145/2993422.2993577
  • Allemang, Dean, and James Hendler. Semantic Web for the Working Ontologist, 2nd ed. Morgan Kaufmann, 2011. (Chs. 1–3)

4.3 Last third: Models

Learning Objectives:

  • Distinguish supervised from unsupervised learning, and regression from classification, and give a civil/environmental engineering example of each
  • Formulate learning as minimizing expected loss, and distinguish training error from generalization error
  • Explain the bias-variance trade-off and how model flexibility drives it
  • Diagnose overfitting and underfitting from training and validation performance
  • Design a defensible model assessment protocol using train/validation/test splits and cross-validation, including the pitfalls of temporal data

References:

  • James, Gareth, Daniela Witten, Trevor Hastie, Robert Tibshirani, and Jonathan Taylor. An Introduction to Statistical Learning with Applications in Python. Springer, 2023. (Ch. 2) Free at https://www.statlearning.com

Learning Objectives:

  • Formulate the ordinary least squares problem and derive its closed-form solution
  • Interpret estimated coefficients, standard errors, and diagnostic plots
  • Connect least squares to maximum likelihood estimation under Gaussian noise
  • Apply ridge and lasso regularization, and explain how each controls model complexity
  • Engineer features (transformations, interactions, encodings of categorical variables) to improve a linear model
  • Fit and validate a regression model on real sensor data (e.g., predicting building energy use from weather)

References:

  • James, Gareth, et al. An Introduction to Statistical Learning with Applications in Python. Springer, 2023. (Chs. 3, 6)

Learning Objectives:

  • Explain why least squares is ill-suited to classification and formulate logistic regression instead
  • Interpret logistic regression coefficients in terms of log-odds and predicted probabilities
  • Visualize and reason about linear decision boundaries
  • Evaluate classifiers using confusion matrices, precision, recall, and ROC curves, and choose metrics appropriate to the application
  • Recognize the effects of class imbalance and strategies to address it
  • Apply classification to an engineering monitoring problem (e.g., damage or fault detection)

References:

  • James, Gareth, et al. An Introduction to Statistical Learning with Applications in Python. Springer, 2023. (Ch. 4)

Learning Objectives:

  • Extend linear models with polynomial terms, step functions, and splines
  • Explain how basis expansions retain the machinery of linear models while capturing nonlinear relationships
  • Describe k-nearest neighbors and the effect of k on the bias-variance trade-off
  • Recognize the curse of dimensionality and its consequences for local methods
  • Select an appropriate level of model flexibility for a given dataset using validation evidence

References:

  • James, Gareth, et al. An Introduction to Statistical Learning with Applications in Python. Springer, 2023. (Ch. 7)

Learning Objectives:

  • Explain how decision trees are grown, and why single trees tend to overfit
  • Describe how bagging and random forests reduce variance through ensembling
  • Describe boosting and contrast it with bagging
  • Interpret variable importance measures and partial dependence, and state their limitations
  • Apply tree ensembles to tabular engineering data and compare their performance against linear baselines

References:

  • James, Gareth, et al. An Introduction to Statistical Learning with Applications in Python. Springer, 2023. (Ch. 8)

Learning Objectives:

  • Formulate principal component analysis as variance maximization and use it for dimensionality reduction and visualization
  • Apply and compare clustering algorithms (k-means, hierarchical, density-based) and assess cluster quality
  • Frame anomaly detection as an unsupervised problem and apply distance-, density-, and reconstruction-based detectors
  • Connect these tools to monitoring applications: novelty detection in structural health monitoring, load pattern discovery in energy data
  • Articulate why validating unsupervised methods is fundamentally harder than validating supervised ones

References:

  • James, Gareth, et al. An Introduction to Statistical Learning with Applications in Python. Springer, 2023. (Ch. 12)

Learning Objectives:

  • Describe the architecture of a multilayer perceptron: layers, weights, activation functions
  • Explain, conceptually, how networks are trained via gradient descent and backpropagation
  • Recognize the practical levers of training: learning rate, batch size, early stopping, and regularization
  • Explain how recurrent architectures and autoregressive formulations model sequential dependence in time-series data
  • Identify when deep learning is (and is not) warranted for engineering problems, considering data volume, interpretability, and physical consistency

References:

  • James, Gareth, et al. An Introduction to Statistical Learning with Applications in Python. Springer, 2023. (Ch. 10)
  • Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. Deep Learning. MIT Press, 2016. (Chs. 6, 10) Free at https://www.deeplearningbook.org

4.4 Coda: Sensors, Data, and Models

Learning Objectives:

  • Trace a complete pipeline from physical phenomenon to actionable prediction for a real case study, identifying every design decision along the way
  • Explain how choices made at the sensing and data layers (sampling rate, resolution, storage, metadata) constrain what models can later achieve
  • Plan a monitoring campaign end-to-end, drawing on lessons learned from computer systems research about instrumentation, logging, and measurement at scale
  • Critique an end-to-end system design for physical consistency, scalability, and maintainability
  • Map the concepts from each third of the course onto your own final project, and identify the weakest link in your pipeline

References:

  • Selected case studies posted on Canvas.