From 319d791fb6f2bf593e95dcbdae8424cf18e19bb2 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Fri, 11 Sep 2026 13:06:38 +0200 Subject: [PATCH 01/24] Vocabulary from motivation to quick reference Some connection to glossary from the "motivations" episode --- content/motivation.md | 47 ++++++------------------- content/quick-reference.md | 70 ++++++++++++++++++++++++++++++++++++++ 2 files changed, 80 insertions(+), 37 deletions(-) diff --git a/content/motivation.md b/content/motivation.md index 7468fe5b..e5c7472e 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -109,29 +109,30 @@ CORRECT ```{list-table} Problems, Solutions and who is affected? :widths: 40 30 30 -* - Problem - - Solution - - Who is affected? +* - **Problem** + - **Solution** + - **Who is affected?** * - Breaking old functionality when adding new features - - End-to-End tests + - {term}`End-to-End tests` - Developers * - Verify installation - - Smoke tests + - {term}`Smoke tests` - Users * - Showing up-to-date example - - End-to-End tests + - {term}`End-to-End tests` - Users * - Improve readability and names - - Unit tests + - {term}`Unit tests` - Developers -* - Refactor and restructure +* - Change things with confidence + that nothing is breaking - All tests - Developers * - Documentation out of date - Executable notebooks and [nbval](https://github.com/computationalmodelling/nbval), - End-to-End tests + {term}`End-to-End tests` - Users ``` @@ -175,34 +176,6 @@ Use the collaborative notes to answer these questions: --- -## Testing vocabulary - -* Test functions and methods one at a time - **Unit tests** - -* Test how parts work together - **Integration tests** - -* Test the whole thing running, checking the output - **End-to-end tests** - * For example, running on sample data. - -* Test that the whole thing runs in the simplest scenario possible - **Smoke Test** - * if this fails, no point in testing other things, usually. - -* Check results are the same as before - **Regression tests** - * Other names for the same thing: **Acceptance Tests**, **Golden-Master Tests**, **Characterization Tests** - -* Write test first (the output), then write code to make test pass - - **Test-driven development** - -* GitHub or GitLab runs tests automatically - **Continuous - integration** - -* Report that tells you which lines were/were not run by tests - - **Code coverage** - -* Framework that runs test for you - **Testing framework** - * See [Quick Reference](./quick-reference) for some examples. - ---- ## What should you do? diff --git a/content/quick-reference.md b/content/quick-reference.md index 3cff9f44..5b5c5da9 100644 --- a/content/quick-reference.md +++ b/content/quick-reference.md @@ -1,5 +1,75 @@ # Quick Reference +## Glossary +:::::{glossary} + +Unit test + A test that covers a single functions or method + (a "unit") + +Integration test + A test that checks if "units" work together as intended + +Smoke Test + Check that the whole application/script runs without errors + in the simplest scenario possible. + If this fails, no point in testing other things, usually. + +End-to-end test + Test the whole thing running, checking the output + (For example, running on sample data + and checking that the output is ) + +Regression tests + Check results are the same as before, typically written + once a *regression* is detected + (a loss of functionality due to a bug). + Other names for the same kind of test: + - Acceptance Tests + - Golden-Master Tests + - Characterization Tests + +Characterization Tests + Test that are written on existing (legacy) code + in order to have a test suite + to make further work easier + +Test-first development + The practice of writing automated tests + before writing the code that makes the tests pass + +Test-driven development + A special case of Test-First development + where the workflow is + + - Make a list of specifications + + Then, for each specification: + - Write a test, and verify that the test fails + - Write the minimum amount of code to make the test pass + - Refactor and improve the code + + +Continuous integration + The practice of merging in the main branch frequently + without having long-lived branches. + + This typically requires automating part of the workflow, + especially testing, + and GitHub/GitLab et similia have support for that + via Actions/CI-CD respectively. + +Code coverage + Metric representing the fraction of your code base + executed during the test suite. + +Testing framework + Framework that runs test for you. + See the following for some examples. + + +::::: + ## Available tools ### Unit test frameworks From 03bf5d9a30faad44826c43a80e37113182981dd8 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Fri, 11 Sep 2026 13:34:45 +0200 Subject: [PATCH 02/24] Fix "what can test help you do?" section removed paragraph whose topic was already covered in the table --- content/motivation.md | 10 ++++------ 1 file changed, 4 insertions(+), 6 deletions(-) diff --git a/content/motivation.md b/content/motivation.md index e5c7472e..8dbbc242 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -137,13 +137,11 @@ CORRECT ``` -**Help other developers modify it** -- Change things with confidence that nothing is breaking. -- Warning if documentation/examples go out of date. +Moreover, **code that is easy to test is probably easier to maintain**, +since it needs to be more modular and have better separation of concerns. -**Manage complexity** -- If code is easy to test, it's probably easier to maintain. -- The next lesson [Modular code development](https://coderefinery.github.io/modular-type-along/) + +The [Modular code development](https://coderefinery.github.io/modular-type-along/) lesson demonstrates this. --- From bdbe948ff02ceedd1f68f94712b57695a8697034 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Fri, 11 Sep 2026 15:15:29 +0200 Subject: [PATCH 03/24] moved recommendations to the conclusions --- content/conclusions.md | 83 +++++++++++++++++++++++++++----------- content/motivation.md | 58 +++++++++++++------------- content/quick-reference.md | 20 ++++++--- 3 files changed, 102 insertions(+), 59 deletions(-) diff --git a/content/conclusions.md b/content/conclusions.md index d2872a74..bc84a3fd 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -1,39 +1,74 @@ # Conclusions and recommendations -## The basics +```{objectives} +- What is a realistic approach to automated test? +- How can I start? +``` -- Learn one test framework well enough for basics - - Explore and use the good tools that exist out there - - An incomplete list of testing frameworks can be found in the [Quick Reference](quick-reference) -- Start with some basics - - Some simple thing that test all parts -- Automate tests - - Faster feedback and reduce the number of surprises +--- +## Discussion: What's easy and hard to test? -## Going more in-depth +```{discussion} Discussion: Testing in practice -- Strike a healthy balance between unit tests and integration tests -- As the code gets larger and the chance of undetected bugs - increases, tests should increase -- When adding new functionality, also add tests -- When you discover and fix a bug, also commit a test against this bug -- Use code coverage analysis to identify untested or unused code -- If you make your code easier to test, it becomes more modular +Use the collaborative notes to answer these questions: + +1. Give examples of things (from your work) that are easy to test. +2. Give examples of things (from your work) that are hard to test. +``` + +--- + +Considering automated tests when writing code +is a major mental shift, that you will hopefully embrace +after attending this lesson. + +## The basics: what knowledge do you need? + +Learn one test framework well enough for basics: +- Explore and use the good tools that exist out there +- An incomplete list of testing frameworks can be found in the [Quick Reference](quick-reference) + + +## Don't over-test +- Not every code needs perfect {term}`test coverage`. +- A simple script or notebook probably does not need an automated test. -## Ways to get started +## Take the low-hanging fruits first You probably won't do everything perfectly when you start off... But what are some of the easy starting points? -- Do you have some single functions that are easy to test, but hard to - verify just by looking at them? Add unit tests. +**If you have not got anything yet**: +1. Start with an end-to-end test. + Typically easy to add, from a typical "manual" use case. + This should match (or serve as) an **example in the code documentation** anyway. + - Describe in words how *you* check whether the code still works. + - Translate the words into a script. + - Run the script as often as reasonable -- Do you have data analysis or simulation of some sort? Make an - end-to-end test with sample data, or sample parameters. This is - useful as an example anyway. +2. Do you have some single functions that are easy to test, but hard to + verify just by looking at them? Add unit tests. -- A local testing framework + GitHub actions is very easy! And works - well in the background - you do whatever you want and get an email +3. A local testing framework + GitHub actions/Gitlab CI-CD is very easy! + And works well in the background - you do whatever you want and get an email if you break things. It's actually pretty freeing. + +**If you need to start modifying some existing code:** +1. Add a {term}`characterization test` for the part of the code you need to change +2. Add tests for any functionality you intend to add +3. Consider adding some end-to-end tests for the use case you have in mind. + +## Going more in-depth + +- Strike a healthy balance between unit tests and integration tests +- As the code gets larger and the chance of undetected bugs + increases, tests should increase +- When adding new functionality, also add tests +- When you discover and fix a bug, also commit a test against this bug +- Use {term}`code coverage` analysis to identify untested or unused code. + Remember [Goodhart's Law](https://en.wikipedia.org/wiki/Goodhart%27s_law) +- If you make your code easier to test, it becomes more modular +- **Learning how to test well make the rest of your code better, too.** + diff --git a/content/motivation.md b/content/motivation.md index 8dbbc242..7f9e44d0 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -16,7 +16,7 @@ How can we avoid problems like these? ## What are typical problems that *automated* tests can address? -Have you ever had some of these problems? +Have you ever had any of these problems? - You change B and C, and suddenly A doesn't work anymore. Time wasted trying to figure out what changed. @@ -114,28 +114,37 @@ CORRECT - **Who is affected?** * - Breaking old functionality when adding new features - - {term}`End-to-End tests` + - {term}`End-to-End test`s - Developers * - Verify installation - - {term}`Smoke tests` + - {term}`Smoke test`s - Users -* - Showing up-to-date example - - {term}`End-to-End tests` - - Users -* - Improve readability and names - - {term}`Unit tests` +* - Make small incremental changes, + Like improving readability, names + - {term}`Unit test`s + - Developers +* - Make architectural changes, + Like shifting code between + classes, modules and functions + - {term}`Integration test`s + {term}`End-to-End test`s - Developers * - Change things with confidence that nothing is breaking - All tests - Developers -* - Documentation out of date +* - Documentation out of date + including code examples - Executable notebooks and [nbval](https://github.com/computationalmodelling/nbval), - {term}`End-to-End tests` + {term}`End-to-End test`s - Users ``` +Very few people are proud of the code they write +the first time they write it. +Code without automated tests cannot be improved as easily +as code with automated tests. Moreover, **code that is easy to test is probably easier to maintain**, since it needs to be more modular and have better separation of concerns. @@ -145,10 +154,9 @@ The [Modular code development](https://coderefinery.github.io/modular-type-along demonstrates this. --- +## Discussion: When is it OK not to add automated tests? -## Discussion: When is it OK not to add tests? - -```{discussion} Discussion: When is it OK not to add tests? +::::{discussion} Discussion: When is it OK not to add automated tests? Vote in the notes and we'll discuss soon. **It is always a balance: there is no "always"/"never"**. @@ -158,26 +166,25 @@ Vote in the notes and we'll discuss soon. **It is always a balance: there is no 3. A simple short, "obviously correct" shell script? 4. Can you give other examples? -``` +:::{solution} +The role of automated tests is to save time when making changes to code. ---- +1. In this case you just "test manually" the notebook by running it. Automated tests might not save you time. + But if some non trivial functions are added, you might want to have automated {term}`unit test`s for these separately. -## Discussion: What's easy and hard to test? +2. Writing automated tests "Throwaway code" can be a waste of time. But if you get back to it, then you should think about writing automated tests for it. -```{discussion} Discussion: Testing in practice +3. "manual test" can be sufficient. In case of changes, checking the script with a {term}`linter` like [Shellcheck](https://www.shellcheck.net/) might still be useful! -Use the collaborative notes to answer these questions: +::: -1. Give examples of things (from your work) that are easy to test. -2. Give examples of things (from your work) that are hard to test. -``` +:::: --- ## What should you do? -* Not every code needs perfect test coverage. * If code is interactive-only (Jupyter Notebook), it's usually hard to test. @@ -202,13 +209,6 @@ Use the collaborative notes to answer these questions: ## Where to start -- A simple script or notebook probably does not need an automated test. - -**If you have nothing yet** -- Start with an end-to-end test. -- Describe in words how *you* check whether the code still works. -- Translate the words into a script. -- Run the script automatically on every code change. **If you want to start with unit-testing** - You want to rewrite a function? Start adding a unit test right there first. diff --git a/content/quick-reference.md b/content/quick-reference.md index 5b5c5da9..0828621f 100644 --- a/content/quick-reference.md +++ b/content/quick-reference.md @@ -13,14 +13,15 @@ Integration test Smoke Test Check that the whole application/script runs without errors in the simplest scenario possible. - If this fails, no point in testing other things, usually. + If this fails, no point in testing other things, usually End-to-end test Test the whole thing running, checking the output (For example, running on sample data - and checking that the output is ) + and checking that the output is the expected one - + see also Regression test) -Regression tests +Regression test Check results are the same as before, typically written once a *regression* is detected (a loss of functionality due to a bug). @@ -29,9 +30,9 @@ Regression tests - Golden-Master Tests - Characterization Tests -Characterization Tests - Test that are written on existing (legacy) code - in order to have a test suite +Characterization Test + Automated test that is written on existing (legacy) code + and added to a test suite to make further work easier Test-first development @@ -62,11 +63,18 @@ Continuous integration Code coverage Metric representing the fraction of your code base executed during the test suite. + Note: this is an easy-to-game metric + and it is perfectly possible to write tests that are completely useless + but increase code coverage. Testing framework Framework that runs test for you. See the following for some examples. +Linter + A program that can check your code + for typical mistakes + or for risky practices, and reports them to you. ::::: From b13270d51c303e0f3d2663c719d8530ed4c2e16d Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Thu, 17 Sep 2026 17:35:53 +0200 Subject: [PATCH 04/24] Add comment on self-hosting Addressing comment in https://github.com/coderefinery/testing/pull/246#pullrequestreview-5171588513 --- content/remotely.md | 17 +++++++++++++++-- 1 file changed, 15 insertions(+), 2 deletions(-) diff --git a/content/remotely.md b/content/remotely.md index f75abf4c..a95554f4 100644 --- a/content/remotely.md +++ b/content/remotely.md @@ -579,15 +579,28 @@ Finally, we discuss together about our experiences with this exercise. ## Where to go from here -- This example was using Python but you can achieve the same automation for R or Fortran or C/C++ or other languages -- This workflow is very useful for collaborators who work on the same code and it works both for +**These techniques:** +- **apply to all programming languages**: + This example was using Python but you can achieve the same automation for R or Fortran or C/C++ or other languages +- **are recommended and useful for collaborative software develpment**: This workflow is very useful for collaborators who work on the same code and it works both for [centralized](https://coderefinery.github.io/git-collaborative/02-centralized/) and [forking](https://coderefinery.github.io/git-collaborative/03-distributed/) workflows - have a look at this [alternative exercise](./full-cycle-ci) to see how that works. + +**There is more tooling available** to discover, for example: + - GitHub Actions has a [Marketplace](https://github.com/marketplace?type=actions) which offer wide range of automatic workflows - On GitLab use [GitLab CI](https://about.gitlab.com/product/continuous-integration/) - For Windows builds you can also use [Appveyor](https://www.appveyor.com) +**About self-hosting**: +- Note that this works also for self-hosted GitLab servers or GitHub Enterprise + (i.e., not gitlab.com or github.com). + Your data does not have to go to a cloud! +- It is also possible to run these workflows on a host + which is not the one where the repository is hosted + (including a HPC system) + if you need more control on the running environment. ```{keypoints} - When fixing bugs or other problems reported in issues, use the issue From afad9c8bdfc6c7d05cb26db564aa7b3370caa1e5 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Sat, 19 Sep 2026 15:21:54 +0200 Subject: [PATCH 05/24] Small rewording fixes Co-Authored-by: Anja Virkkunen Co-authored-by: Michele Mesiti --- content/conclusions.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/content/conclusions.md b/content/conclusions.md index bc84a3fd..1f0972ce 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -35,12 +35,12 @@ Learn one test framework well enough for basics: - Not every code needs perfect {term}`test coverage`. - A simple script or notebook probably does not need an automated test. -## Take the low-hanging fruits first +## Pick the low-hanging fruits first You probably won't do everything perfectly when you start off... But what are some of the easy starting points? -**If you have not got anything yet**: +**If you have got nothing yet**: 1. Start with an end-to-end test. Typically easy to add, from a typical "manual" use case. This should match (or serve as) an **example in the code documentation** anyway. From e68959f398f74b5634830e95354030d9f98e6ece Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Sat, 19 Sep 2026 15:27:48 +0200 Subject: [PATCH 06/24] Completely remove "where to start" from motivation it belongs to conclusions, where the topic is already covered. --- content/motivation.md | 6 ------ 1 file changed, 6 deletions(-) diff --git a/content/motivation.md b/content/motivation.md index 7f9e44d0..d2fc1f17 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -205,10 +205,4 @@ The role of automated tests is to save time when making changes to code. * **Learning how to test well make the rest of your code better, too.** ---- - -## Where to start - -**If you want to start with unit-testing** -- You want to rewrite a function? Start adding a unit test right there first. From 4fd8d0d6111b32f3b26967e0ac34d9ca0a67d27f Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Sat, 19 Sep 2026 18:13:12 +0200 Subject: [PATCH 07/24] Elaborate on runner self-hosting Folliwing up from https://github.com/coderefinery/testing/pull/252#discussion_r4046437914 --- content/remotely.md | 31 ++++++++++++++++++++++--------- 1 file changed, 22 insertions(+), 9 deletions(-) diff --git a/content/remotely.md b/content/remotely.md index a95554f4..8f4654bb 100644 --- a/content/remotely.md +++ b/content/remotely.md @@ -589,18 +589,31 @@ Finally, we discuss together about our experiences with this exercise. **There is more tooling available** to discover, for example: -- GitHub Actions has a [Marketplace](https://github.com/marketplace?type=actions) which offer wide range of automatic workflows -- On GitLab use [GitLab CI](https://about.gitlab.com/product/continuous-integration/) -- For Windows builds you can also use [Appveyor](https://www.appveyor.com) +- GitHub Actions: + - has a [Marketplace](https://github.com/marketplace?type=actions) which offer wide range of composable actions + - provides template *workflows* to start from + - Can execute *workflows* on on various [runners](https://docs.github.com/en/actions/concepts/runners/github-hosted-runners) +- On GitLab the equivalent feature is named [GitLab CI/CD](https://about.gitlab.com/product/continuous-integration/) + - composability is achieved using so-called ["components"](https://gitlab.com/explore/catalog). + - Examples of typical *pipelines* are also [available](https://docs.gitlab.com/ci/examples/). + - Can execute *pipelines* on on various [runners](https://docs.gitlab.com/ci/runners/#gitlab-hosted-runners) -**About self-hosting**: -- Note that this works also for self-hosted GitLab servers or GitHub Enterprise +Both github.com and gitlab.com offer runners on Linux, Windows and MacOS machines. + + +If you need more control **About self-hosting**: +- CI/CD and Actions are (typically) also available + for self-hosted GitLab servers or GitHub Enterprise (i.e., not gitlab.com or github.com). Your data does not have to go to a cloud! -- It is also possible to run these workflows on a host - which is not the one where the repository is hosted - (including a HPC system) - if you need more control on the running environment. +- If you need more control on the running environment, + you can also connect the git server + to another host + (including, e.g., a HPC system) + and run workflows/pipelines there, + with self-hosted runners + (for [GitHub](https://docs.github.com/en/actions/concepts/runners/self-hosted-runners) + and for [GitLab](https://docs.gitlab.com/runner/)). ```{keypoints} - When fixing bugs or other problems reported in issues, use the issue From 17a80d21b27b851c8d00f583a5f0b8acfd66fdbc Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Sat, 19 Sep 2026 18:20:41 +0200 Subject: [PATCH 08/24] Motivations: small reformatting for readability Follow up from https://github.com/coderefinery/testing/pull/252#discussion_r4046566738 --- content/motivation.md | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/content/motivation.md b/content/motivation.md index d2fc1f17..e1b251cb 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -5,8 +5,7 @@ - Understand various benefits of testing ``` -Most scientists nowadays depend on software for research. - +Most scientists nowadays depend on software for research. What can go wrong when research software has bugs? Look no further: - [A Scientist's Nightmare: Software Problem Leads to Five Retractions](https://science.sciencemag.org/content/314/5807/1856.summary) From 52aae8599c2e658d37d79f3e349ca48d1dd45349 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Sat, 19 Sep 2026 18:26:01 +0200 Subject: [PATCH 09/24] fix typo --- content/remotely.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/content/remotely.md b/content/remotely.md index 8f4654bb..364ae9a0 100644 --- a/content/remotely.md +++ b/content/remotely.md @@ -582,7 +582,7 @@ Finally, we discuss together about our experiences with this exercise. **These techniques:** - **apply to all programming languages**: This example was using Python but you can achieve the same automation for R or Fortran or C/C++ or other languages -- **are recommended and useful for collaborative software develpment**: This workflow is very useful for collaborators who work on the same code and it works both for +- **are recommended and useful for collaborative software development**: This workflow is very useful for collaborators who work on the same code and it works both for [centralized](https://coderefinery.github.io/git-collaborative/02-centralized/) and [forking](https://coderefinery.github.io/git-collaborative/03-distributed/) workflows - have a look at this [alternative exercise](./full-cycle-ci) to see how that works. From 5cfed987465bd6fff2772aaa75be27f24d4a8462 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Sat, 19 Sep 2026 18:32:50 +0200 Subject: [PATCH 10/24] Use epigraph directive for quote in motivations Keeping the italics, though. As suggested here: https://github.com/coderefinery/testing/pull/252#discussion_r4046609163 Co-Authored-by: Anja Virkkunen --- content/motivation.md | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/content/motivation.md b/content/motivation.md index e1b251cb..377503cf 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -39,12 +39,14 @@ places it's useful for research code, and how easy it can be. ## Untested software can be compared to uncalibrated measurement devices -*"Before relying on a new experimental device, an experimental scientist always +```{epigraph} +*Before relying on a new experimental device, an experimental scientist always establishes its accuracy. A new detector is calibrated when the scientist observes its responses to known input signals. The results of this -calibration are compared against the expected response."* +calibration are compared against the expected response.* -> [From [Testing and Continuous Integration with Python](https://carpentries-incubator.github.io/python-testing/), created by K. Huff] +-- From [Testing and Continuous Integration with Python](https://carpentries-incubator.github.io/python-testing/), created by K. Huff +``` With testing, simulations and analysis using software *can* be held to the same standards as experimental measurement devices! From 598f314f57a488fa16e54abb4e16bde324c39af5 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Sat, 19 Sep 2026 19:50:19 +0200 Subject: [PATCH 11/24] Motivations: refactor + frameworks --- content/motivation.md | 163 +++++++++++++++++++++++++------------ content/quick-reference.md | 65 ++++++++++----- content/remotely.md | 2 +- 3 files changed, 155 insertions(+), 75 deletions(-) diff --git a/content/motivation.md b/content/motivation.md index 377503cf..0dd9f7a9 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -52,12 +52,73 @@ With testing, simulations and analysis using software *can* be held to the same --- +## What can tests help you do? + +```{list-table} Problems, Solutions and who is affected? +:widths: 40 30 30 +* - **Problem** + - **Solution** + - **Who is affected?** +* - Breaking old functionality + when adding new features + - {term}`End-to-End test`s + - Developers +* - Verify installation + - {term}`Smoke test`s + - Users +* - Make small incremental changes, + Like improving readability, names + - {term}`Unit test`s + - Developers +* - Make architectural changes, + Like shifting code between + classes, modules and functions + - {term}`Integration test`s + {term}`End-to-End test`s + - Developers +* - Change things with confidence + that nothing is breaking + - All tests + - Developers +* - Documentation out of date + including code examples + - Executable notebooks + and [nbval](https://github.com/computationalmodelling/nbval), + {term}`End-to-End test`s + - Users + +``` +Very few people are proud of the code they write +the first time they write it. +Code without automated tests cannot be improved as easily +as code with automated tests. + +Moreover, **code that is easy to test is probably easier to maintain**, +since it needs to be more modular and have better separation of concerns. + + +The [Modular code development](https://coderefinery.github.io/modular-type-along/) lesson + demonstrates this. + +--- + ## Testing in a nutshell -In the most basic form of software tests, -expected results are compared with observed results -in order to establish accuracy. Why are we not comparing directly all -digits with the expected result? +There are many forms of testing. + +One can write test programs and run them: +```console +$ python3 run-test.py + +running: sample_data/set1.csv --output=tests/set1.txt +CORRECT +``` + +In the most basic form of a software test, +the observed result is compared with expected result (an "*oracle*") +in order to establish correctness. +Here are some examples of this testing pattern +in different programming languages: ````{tabs} ```{group-tab} Python @@ -96,65 +157,60 @@ digits with the expected result? ``` ```` -Or you can test whole programs: -```console -$ python3 run-test.py +### Why use a testing framework? -running: sample_data/set1.csv --output=tests/set1.txt -CORRECT -``` +Automated testing typically requires to do a number of repetitive tasks +and to solve some tricky problems. ---- +Fortunately for us, +someone has already found a solution for most of these +and created {term}`testing framework`s that we can use +(see [Unit test frameworks](./quick-reference.md#unit-test-frameworks)). -## What can tests help you do? +Note: not all frameworks solve all problems +(also because sometimes the underlying language does not have the necessary features). -```{list-table} Problems, Solutions and who is affected? -:widths: 40 30 30 +```{list-table} Why use a testing framework? * - **Problem** - **Solution** - - **Who is affected?** -* - Breaking old functionality - when adding new features - - {term}`End-to-End test`s - - Developers -* - Verify installation - - {term}`Smoke test`s - - Users -* - Make small incremental changes, - Like improving readability, names - - {term}`Unit test`s - - Developers -* - Make architectural changes, - Like shifting code between - classes, modules and functions - - {term}`Integration test`s - {term}`End-to-End test`s - - Developers -* - Change things with confidence - that nothing is breaking - - All tests - - Developers -* - Documentation out of date - including code examples - - Executable notebooks - and [nbval](https://github.com/computationalmodelling/nbval), - {term}`End-to-End test`s - - Users - +* - Rememer to run all the test you write + in the main test script/program + - automatic discovery, + automatic registration + when declaring test functions +* - Run only some tests + (to save time) + - Test filtering + via patterns +* - Report failures/successes + (to humans or other machines) + - Automated collection and output, + (e.g., Junit XML format or [TAP](https://en.wikipedia.org/wiki/Test_Anything_Protocol)) +* - Provide useful information + on why a test has failed + - "smart" assertions/macros +* - Run same test for many + known input/output combinations + - Parametric tests +* - Check that a property holds + for a class of inputs and outputs + - Automatically generate test cases + based on a strategy + (property testing) +* - Debugging on failure + - Start debugger on test failure +* - Floating point equalities with tolerance + - macros/classes +* - Set up and tear down complex test cases + - {term}`Fixture`s +* - Estimate how much of your code + is *run* in the test suite + - automatic {term}`code coverage` measurement ``` -Very few people are proud of the code they write -the first time they write it. -Code without automated tests cannot be improved as easily -as code with automated tests. -Moreover, **code that is easy to test is probably easier to maintain**, -since it needs to be more modular and have better separation of concerns. - - -The [Modular code development](https://coderefinery.github.io/modular-type-along/) lesson - demonstrates this. --- + ## Discussion: When is it OK not to add automated tests? ::::{discussion} Discussion: When is it OK not to add automated tests? @@ -183,7 +239,6 @@ The role of automated tests is to save time when making changes to code. --- - ## What should you do? diff --git a/content/quick-reference.md b/content/quick-reference.md index 0828621f..e86a7080 100644 --- a/content/quick-reference.md +++ b/content/quick-reference.md @@ -11,37 +11,46 @@ Integration test A test that checks if "units" work together as intended Smoke Test - Check that the whole application/script runs without errors + Check that the whole application or script runs without errors in the simplest scenario possible. If this fails, no point in testing other things, usually + +Regression + A loss of functionality, typically due to a bug + End-to-end test - Test the whole thing running, checking the output + Test the whole thing running, checking the output (For example, running on sample data - and checking that the output is the expected one - - see also Regression test) + and checking that the output is the expected one. + See also {term}`Regression test`) + Regression test - Check results are the same as before, typically written - once a *regression* is detected - (a loss of functionality due to a bug). + + Check that results and behaviour are what they are supposed to be. + Typically written once a {term}`regression` is detected. Other names for the same kind of test: + - Characterization Tests - Acceptance Tests - Golden-Master Tests - - Characterization Tests - + Characterization Test - Automated test that is written on existing (legacy) code + + Automated test that is written on existing (legacy) code, + (assuming that version of the code is correct) and added to a test suite - to make further work easier + to make further work/changes easier Test-first development + The practice of writing automated tests before writing the code that makes the tests pass Test-driven development + A special case of Test-First development - where the workflow is + where the workflow is: - Make a list of specifications @@ -52,30 +61,48 @@ Test-driven development Continuous integration + The practice of merging in the main branch frequently without having long-lived branches. - This typically requires automating part of the workflow, especially testing, and GitHub/GitLab et similia have support for that via Actions/CI-CD respectively. Code coverage - Metric representing the fraction of your code base - executed during the test suite. - Note: this is an easy-to-game metric + + Metric representing the fraction of code base + executed during the test suite. + Note: this is only an **upper bound** + to the fraction of code base + that is anyhow tested. and it is perfectly possible to write tests that are completely useless but increase code coverage. Testing framework + Framework that runs test for you. See the following for some examples. Linter + A program that can check your code for typical mistakes or for risky practices, and reports them to you. +Fixture + A resource that needs to be set up before a test case can run + and needs to be torn down after the test case + (or a whole test suite) + has run. + +Property testing + test that a property of the code holds + for a whole class of inputs. + Typically done by automatically generating + Typically very time-consuming + compared to unit testing. + ::::: ## Available tools @@ -359,10 +386,8 @@ You can then compile using this script: Each of these are web services to handle testing, free for open source projects. -- [GitHub Actions](https://github.com/features/actions) (we will - demonstrate this in the next episode) -- [GitLab CI](https://about.gitlab.com/features/gitlab-ci-cd/) - (we will demonstrate this in the next episode) +- [GitHub Actions](https://github.com/features/actions) - see episode [Automated Testing Remotely](./remotely) +- [GitLab CI](https://about.gitlab.com/features/gitlab-ci-cd/) - see episode [Automated Testing Remotely](./remotely) - [Azure Pipelines](https://azure.microsoft.com/en-us/services/devops/pipelines/) - [Coveralls](https://coveralls.io) - [Codecov](https://codecov.io) diff --git a/content/remotely.md b/content/remotely.md index 364ae9a0..a8ab4b40 100644 --- a/content/remotely.md +++ b/content/remotely.md @@ -27,7 +27,7 @@ In this exercise, we will: - **C.** Find a bug in our repository and open an issue to report it - **D.** Fix the bug on a bugfix branch and open a pull request (GitHub)/ merge request (GitLab) - **E.** Merge the pull/merge request and see how the issue is automatically closed. -- **F.** Create a test to increase the code coverage of our tests. +- **F.** Create a test to increase the {term}`code coverage` of our tests. ``` ## Prerequisites From f5d40628f283d6651dbcfd7eac3b805e2ec20090 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Tue, 22 Sep 2026 18:39:09 +0200 Subject: [PATCH 12/24] Fixes from Anja's review Co-Authored-by: Anja Virkkunen --- content/motivation.md | 57 ++++++++++++++++++++++++++++---------- content/quick-reference.md | 20 ++++++++++--- content/remotely.md | 53 +++++++++++++++++++++++++---------- 3 files changed, 96 insertions(+), 34 deletions(-) diff --git a/content/motivation.md b/content/motivation.md index 0dd9f7a9..aa850ef8 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -54,7 +54,7 @@ With testing, simulations and analysis using software *can* be held to the same ## What can tests help you do? -```{list-table} Problems, Solutions and who is affected? +```{list-table} Problems, Solution and who is affected? :widths: 40 30 30 * - **Problem** - **Solution** @@ -171,41 +171,68 @@ Note: not all frameworks solve all problems (also because sometimes the underlying language does not have the necessary features). ```{list-table} Why use a testing framework? +:widths: 40 40 20 * - **Problem** - **Solution** -* - Rememer to run all the test you write + - **Examples** +* - Report failures/successes + (to humans or other machines) + - Automated collection and output, + (e.g., Junit XML format or [TAP](https://en.wikipedia.org/wiki/Test_Anything_Protocol)) + - Fundamental feature + all testing frameworks + have it. +* - Remember to run all the test you write in the main test script/program - - automatic discovery, - automatic registration + - Automatic discovery, + automatic {term}`registration` when declaring test functions + - [Pytest](https://docs.pytest.org/en/stable/) * - Run only some tests (to save time) - Test filtering via patterns -* - Report failures/successes - (to humans or other machines) - - Automated collection and output, - (e.g., Junit XML format or [TAP](https://en.wikipedia.org/wiki/Test_Anything_Protocol)) + - * - Provide useful information on why a test has failed - - "smart" assertions/macros + - "Smart" assertions/macros + - [in pytest](https://docs.pytest.org/en/stable/how-to/assert.html#assert), + [in GoogleTest](https://google.github.io/googletest/primer.html#assertions) * - Run same test for many known input/output combinations - Parametric tests + - [In Pytest](https://docs.pytest.org/en/stable/how-to/parametrize.html#pytest-mark-parametrize-parametrizing-test-functions), + [in Julia](https://docs.julialang.org/en/v1/stdlib/Test/#Working-with-Test-Sets) (see `testset for`) * - Check that a property holds for a class of inputs and outputs - Automatically generate test cases based on a strategy (property testing) + - [hypothesis](https://hypothesis.readthedocs.io/en/latest/tutorial/introduction.html) + (python) * - Debugging on failure - Start debugger on test failure + - `pytest --pdb` * - Floating point equalities with tolerance - - macros/classes + - Macros/classes + - `≈` (julia), `pytest.approx` * - Set up and tear down complex test cases - {term}`Fixture`s + - [In pytest](https://docs.pytest.org/en/stable/explanation/fixtures.html), + [in GoogleTest](https://google.github.io/googletest/primer.html#same-data-multiple-tests) * - Estimate how much of your code is *run* in the test suite - - automatic {term}`code coverage` measurement + - Automatic {term}`coverage` + measurement + - [Pytest-cov](https://pytest-cov.readthedocs.io/en/latest/), + [gcov/lcov](https://wiki.cs.jmu.edu/reference/gcov/) + (for C/C++/Fortran) +* - Will my code work + with different versions + of the dependencies? + - Test in different environments + - [Tox](https://tox.wiki/en), [nox](https://nox.thea.codes/en/stable/index.html) + (python) ``` @@ -229,9 +256,9 @@ The role of automated tests is to save time when making changes to code. 1. In this case you just "test manually" the notebook by running it. Automated tests might not save you time. But if some non trivial functions are added, you might want to have automated {term}`unit test`s for these separately. -2. Writing automated tests "Throwaway code" can be a waste of time. But if you get back to it, then you should think about writing automated tests for it. +2. Writing automated tests for "throwaway code" can be a waste of time. But if you get back to it, then you should think about writing automated tests for it. -3. "manual test" can be sufficient. In case of changes, checking the script with a {term}`linter` like [Shellcheck](https://www.shellcheck.net/) might still be useful! +3. "Manual test" can be sufficient. In case of changes, checking the script with a {term}`linter` like [Shellcheck](https://www.shellcheck.net/) might still be useful! ::: @@ -245,7 +272,7 @@ The role of automated tests is to save time when making changes to code. * If code is interactive-only (Jupyter Notebook), it's usually hard to test. - * But also hard to run: the next lesson will discuss! + * But also hard to run: the [next lesson](https://coderefinery.github.io/modular-type-along/) will discuss! * At least end-to-end is often easy to add. @@ -257,7 +284,7 @@ The role of automated tests is to save time when making changes to code. * It's easy to have Gitlab/Github run the tests. * It's nice to push without thinking, and the system tells you when - it's broke. + it's broken. * **Learning how to test well make the rest of your code better, too.** diff --git a/content/quick-reference.md b/content/quick-reference.md index e86a7080..13226078 100644 --- a/content/quick-reference.md +++ b/content/quick-reference.md @@ -47,7 +47,8 @@ Test-first development The practice of writing automated tests before writing the code that makes the tests pass -Test-driven development +TDD + Test-driven development A special case of Test-First development where the workflow is: @@ -76,12 +77,12 @@ Code coverage Note: this is only an **upper bound** to the fraction of code base that is anyhow tested. - and it is perfectly possible to write tests that are completely useless + And it is perfectly possible to write tests that are completely useless but increase code coverage. Testing framework - Framework that runs test for you. + Framework that runs tests for you. See the following for some examples. Linter @@ -100,8 +101,19 @@ Property testing test that a property of the code holds for a whole class of inputs. Typically done by automatically generating - Typically very time-consuming + test cases according to a strategy. + Tends to very time-consuming compared to unit testing. + +Test Registration + The act of marking a test for execution + in a main testing program. + Test frameworks allow to do this + automatically + at test definition + so that it does not have to be manually invoked + in the "main" script/function, + with the risk of forgetting it. ::::: diff --git a/content/remotely.md b/content/remotely.md index a8ab4b40..eed6ea75 100644 --- a/content/remotely.md +++ b/content/remotely.md @@ -581,30 +581,53 @@ Finally, we discuss together about our experiences with this exercise. **These techniques:** - **apply to all programming languages**: - This example was using Python but you can achieve the same automation for R or Fortran or C/C++ or other languages -- **are recommended and useful for collaborative software development**: This workflow is very useful for collaborators who work on the same code and it works both for + This example was using Python but you can achieve the same automation for any other languages +- **are recommended and useful for collaborative software development**: + automatically running the test suite remotely + as presented here + is very useful for collaborators who work on the same code and it works both for [centralized](https://coderefinery.github.io/git-collaborative/02-centralized/) and [forking](https://coderefinery.github.io/git-collaborative/03-distributed/) workflows - have a look at this [alternative exercise](./full-cycle-ci) to see how that works. **There is more tooling available** to discover, for example: - -- GitHub Actions: - - has a [Marketplace](https://github.com/marketplace?type=actions) which offer wide range of composable actions - - provides template *workflows* to start from - - Can execute *workflows* on on various [runners](https://docs.github.com/en/actions/concepts/runners/github-hosted-runners) -- On GitLab the equivalent feature is named [GitLab CI/CD](https://about.gitlab.com/product/continuous-integration/) - - composability is achieved using so-called ["components"](https://gitlab.com/explore/catalog). - - Examples of typical *pipelines* are also [available](https://docs.gitlab.com/ci/examples/). - - Can execute *pipelines* on on various [runners](https://docs.gitlab.com/ci/runners/#gitlab-hosted-runners) + + +```{list-table} A jungle of tools +* - **Feature** + - **GitHub** + **Actions** + - **GitLab** + **CI/CD** +* - Composable actions + - [GitHub Marketplace](https://github.com/marketplace?type=actions), + very mature ecosystem + - [components](https://gitlab.com/explore/catalog) +* - Defaults + - very minimal + (checkout requires an action) + - typical use cases + are "baked in" +* - Examples and templates + - Template workflows + - Some [Pipeline examples](https://docs.gitlab.com/ci/examples/) +* - Execution + - [GitHub-hosted Runners](https://docs.github.com/en/actions/concepts/runners/github-hosted-runners) + - [GitLab-hosted Runners](https://docs.gitlab.com/ci/runners/#gitlab-hosted-runners) +* - Self-hosted + execution + - [Self-hosted runners](https://docs.github.com/en/actions/how-tos/manage-runners/self-hosted-runners) + - [GitLab runner](https://docs.gitlab.com/runner/) + (more customizable) +``` -Both github.com and gitlab.com offer runners on Linux, Windows and MacOS machines. +Both GitHub Actions (on *github.com*) and GitLab CI/CD (on *gitlab.com*) offer runners on Linux, Windows and MacOS machines. -If you need more control **About self-hosting**: -- CI/CD and Actions are (typically) also available +If you need more control, you can choose to **self-host**: +- CI/CD and Actions can also be made available for self-hosted GitLab servers or GitHub Enterprise - (i.e., not gitlab.com or github.com). + (i.e., outside *gitlab.com* or *github.com* domains). Your data does not have to go to a cloud! - If you need more control on the running environment, you can also connect the git server From 6a2db30ada759cd45378f8d0d35d286f4e66abd7 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Tue, 22 Sep 2026 19:02:42 +0200 Subject: [PATCH 13/24] move "why use a testing framework" to conclusions I think it fits better there, after we have shown why automated tests are important. Also before showing automated tests many of the "typical problems" might sound abstruse or unfamiliar. --- content/conclusions.md | 79 +++++++++++++++++++++++++++++++++++ content/motivation.md | 93 ++++-------------------------------------- 2 files changed, 88 insertions(+), 84 deletions(-) diff --git a/content/conclusions.md b/content/conclusions.md index 1f0972ce..d8f9eb76 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -29,6 +29,85 @@ Learn one test framework well enough for basics: - Explore and use the good tools that exist out there - An incomplete list of testing frameworks can be found in the [Quick Reference](quick-reference) +### Why use a testing framework? + +Automated testing typically requires to do a number of repetitive tasks +and to solve some tricky problems. + +Fortunately for us, +someone has already found a solution for most of these +and created {term}`testing framework`s that we can use +(see [Unit test frameworks](./quick-reference.md#unit-test-frameworks)). + +Note: not all frameworks solve all problems +(also because sometimes the underlying language does not have the necessary features). + +```{list-table} Why use a testing framework? +:widths: 40 40 20 +* - **Typical problem** + - **Solution** + - **Examples** +* - Report failures/successes + (to humans or other machines) + consistently + - Automated collection and output, + (e.g., Junit XML format or [TAP](https://en.wikipedia.org/wiki/Test_Anything_Protocol)) + - Fundamental feature + all testing frameworks + have it. +* - Remember to run all the test you write + in the main test script/program + - Automatic discovery, + automatic {term}`registration` + when declaring test functions + - [Pytest](https://docs.pytest.org/en/stable/) +* - Run only some tests + (to save time) + - Test filtering + via patterns + - +* - Provide useful information + on why a test has failed + - "Smart" assertions/macros + - [in pytest](https://docs.pytest.org/en/stable/how-to/assert.html#assert), + [in GoogleTest](https://google.github.io/googletest/primer.html#assertions) +* - Run same test for many + known input/output combinations + - Parametric tests + - [In Pytest](https://docs.pytest.org/en/stable/how-to/parametrize.html#pytest-mark-parametrize-parametrizing-test-functions), + [in Julia](https://docs.julialang.org/en/v1/stdlib/Test/#Working-with-Test-Sets) (see `testset for`) +* - Check that a property holds + for a class of inputs and outputs + - Automatically generate test cases + based on a strategy + (property testing) + - [hypothesis](https://hypothesis.readthedocs.io/en/latest/tutorial/introduction.html) + (python) +* - Debugging on failure + - Start debugger on test failure + - `pytest --pdb` +* - Floating point equalities with tolerance + - Macros/classes + - `≈` (julia), `pytest.approx` +* - Set up and tear down complex test cases + - {term}`Fixture`s + - [In pytest](https://docs.pytest.org/en/stable/explanation/fixtures.html), + [in GoogleTest](https://google.github.io/googletest/primer.html#same-data-multiple-tests) +* - Estimate how much of your code + is *run* in the test suite + - Automatic {term}`coverage` + measurement + - [Pytest-cov](https://pytest-cov.readthedocs.io/en/latest/), + [gcov/lcov](https://wiki.cs.jmu.edu/reference/gcov/) + (for C/C++/Fortran) +* - Will my code work + with different versions + of the dependencies? + - Test in different environments + - [Tox](https://tox.wiki/en), [nox](https://nox.thea.codes/en/stable/index.html) + (python) +``` + ## Don't over-test diff --git a/content/motivation.md b/content/motivation.md index aa850ef8..84871c48 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -67,12 +67,12 @@ With testing, simulations and analysis using software *can* be held to the same - {term}`Smoke test`s - Users * - Make small incremental changes, - Like improving readability, names + (e.g., improving readability, names) - {term}`Unit test`s - Developers * - Make architectural changes, - Like shifting code between - classes, modules and functions + (e.g., shifting code between + classes, modules and functions) - {term}`Integration test`s {term}`End-to-End test`s - Developers @@ -88,16 +88,19 @@ With testing, simulations and analysis using software *can* be held to the same - Users ``` + Very few people are proud of the code they write the first time they write it. -Code without automated tests cannot be improved as easily + +Typically, they'd like to improve it. + +But code without automated tests cannot be improved as easily as code with automated tests. Moreover, **code that is easy to test is probably easier to maintain**, since it needs to be more modular and have better separation of concerns. - -The [Modular code development](https://coderefinery.github.io/modular-type-along/) lesson +The [Modular code development lesson](https://coderefinery.github.io/modular-type-along/) demonstrates this. --- @@ -157,84 +160,6 @@ in different programming languages: ``` ```` -### Why use a testing framework? - -Automated testing typically requires to do a number of repetitive tasks -and to solve some tricky problems. - -Fortunately for us, -someone has already found a solution for most of these -and created {term}`testing framework`s that we can use -(see [Unit test frameworks](./quick-reference.md#unit-test-frameworks)). - -Note: not all frameworks solve all problems -(also because sometimes the underlying language does not have the necessary features). - -```{list-table} Why use a testing framework? -:widths: 40 40 20 -* - **Problem** - - **Solution** - - **Examples** -* - Report failures/successes - (to humans or other machines) - - Automated collection and output, - (e.g., Junit XML format or [TAP](https://en.wikipedia.org/wiki/Test_Anything_Protocol)) - - Fundamental feature - all testing frameworks - have it. -* - Remember to run all the test you write - in the main test script/program - - Automatic discovery, - automatic {term}`registration` - when declaring test functions - - [Pytest](https://docs.pytest.org/en/stable/) -* - Run only some tests - (to save time) - - Test filtering - via patterns - - -* - Provide useful information - on why a test has failed - - "Smart" assertions/macros - - [in pytest](https://docs.pytest.org/en/stable/how-to/assert.html#assert), - [in GoogleTest](https://google.github.io/googletest/primer.html#assertions) -* - Run same test for many - known input/output combinations - - Parametric tests - - [In Pytest](https://docs.pytest.org/en/stable/how-to/parametrize.html#pytest-mark-parametrize-parametrizing-test-functions), - [in Julia](https://docs.julialang.org/en/v1/stdlib/Test/#Working-with-Test-Sets) (see `testset for`) -* - Check that a property holds - for a class of inputs and outputs - - Automatically generate test cases - based on a strategy - (property testing) - - [hypothesis](https://hypothesis.readthedocs.io/en/latest/tutorial/introduction.html) - (python) -* - Debugging on failure - - Start debugger on test failure - - `pytest --pdb` -* - Floating point equalities with tolerance - - Macros/classes - - `≈` (julia), `pytest.approx` -* - Set up and tear down complex test cases - - {term}`Fixture`s - - [In pytest](https://docs.pytest.org/en/stable/explanation/fixtures.html), - [in GoogleTest](https://google.github.io/googletest/primer.html#same-data-multiple-tests) -* - Estimate how much of your code - is *run* in the test suite - - Automatic {term}`coverage` - measurement - - [Pytest-cov](https://pytest-cov.readthedocs.io/en/latest/), - [gcov/lcov](https://wiki.cs.jmu.edu/reference/gcov/) - (for C/C++/Fortran) -* - Will my code work - with different versions - of the dependencies? - - Test in different environments - - [Tox](https://tox.wiki/en), [nox](https://nox.thea.codes/en/stable/index.html) - (python) -``` - --- From 5a068cfc4e441687d70077e68d56a1a8b03a714e Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Tue, 22 Sep 2026 19:25:07 +0200 Subject: [PATCH 14/24] mention e2e and unit test in examples in motivations Co-Authored-by: Anja Virkkunen --- content/motivation.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/content/motivation.md b/content/motivation.md index 84871c48..b13ae3e6 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -109,7 +109,7 @@ The [Modular code development lesson](https://coderefinery.github.io/modular-typ There are many forms of testing. -One can write test programs and run them: +One can write test programs and run them (a form of {term}`end-to-end test`): ```console $ python3 run-test.py @@ -121,7 +121,7 @@ In the most basic form of a software test, the observed result is compared with expected result (an "*oracle*") in order to establish correctness. Here are some examples of this testing pattern -in different programming languages: +in different programming languages (in this case, {term}`unit test`s): ````{tabs} ```{group-tab} Python From 31895164f0a5f6b1f47c75c681cc73a303bdefa0 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Tue, 22 Sep 2026 19:25:57 +0200 Subject: [PATCH 15/24] Some addition to conclusions and quick reference --- content/conclusions.md | 11 +++++++++-- content/quick-reference.md | 13 +++++++------ 2 files changed, 16 insertions(+), 8 deletions(-) diff --git a/content/conclusions.md b/content/conclusions.md index d8f9eb76..98b83d0d 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -27,7 +27,9 @@ after attending this lesson. Learn one test framework well enough for basics: - Explore and use the good tools that exist out there -- An incomplete list of testing frameworks can be found in the [Quick Reference](quick-reference) +- An incomplete list of {term}`testing framework`s + can be found in the [Quick Reference](quick-reference) + (see [Unit test frameworks](./quick-reference.md#unit-test-frameworks)). ### Why use a testing framework? @@ -37,7 +39,6 @@ and to solve some tricky problems. Fortunately for us, someone has already found a solution for most of these and created {term}`testing framework`s that we can use -(see [Unit test frameworks](./quick-reference.md#unit-test-frameworks)). Note: not all frameworks solve all problems (also because sometimes the underlying language does not have the necessary features). @@ -100,6 +101,12 @@ Note: not all frameworks solve all problems - [Pytest-cov](https://pytest-cov.readthedocs.io/en/latest/), [gcov/lcov](https://wiki.cs.jmu.edu/reference/gcov/) (for C/C++/Fortran) +* - Do code examples in documentation + work as expected? + - doctests + - [Python](https://docs.python.org/3/library/doctest.html), + [Julia](https://documenter.juliadocs.org/stable/man/doctests/), + [R](https://cran.r-project.org/web/packages/doctest/vignettes/doctest.html) * - Will my code work with different versions of the dependencies? diff --git a/content/quick-reference.md b/content/quick-reference.md index 13226078..d663d6fb 100644 --- a/content/quick-reference.md +++ b/content/quick-reference.md @@ -48,17 +48,18 @@ Test-first development before writing the code that makes the tests pass TDD - Test-driven development + Acronym for {term}`Test-driven development` +Test-driven development A special case of Test-First development where the workflow is: - - Make a list of specifications + - Make a list of specifications - Then, for each specification: - - Write a test, and verify that the test fails - - Write the minimum amount of code to make the test pass - - Refactor and improve the code + - Then, for each specification: + - Write a test, run it and verify that the test fails + - Write the minimum amount of code to make the test pass + - Refactor and improve the code Continuous integration From 04abe9a61d5983cb2fb0474952ee1141c25a6d1f Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Tue, 22 Sep 2026 19:30:42 +0200 Subject: [PATCH 16/24] modularity <-> testability in conclusions Co-Authored-by: Anja Virkkunen --- content/conclusions.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/content/conclusions.md b/content/conclusions.md index 98b83d0d..7e24aff8 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -155,6 +155,7 @@ what are some of the easy starting points? - When you discover and fix a bug, also commit a test against this bug - Use {term}`code coverage` analysis to identify untested or unused code. Remember [Goodhart's Law](https://en.wikipedia.org/wiki/Goodhart%27s_law) -- If you make your code easier to test, it becomes more modular +- If you make your code easier to test, it becomes more modular (and vice versa - see the [modular code development lesson](https://coderefinery.github.io/modular-type-along/)) + - **Learning how to test well make the rest of your code better, too.** From d48196275eaa57268049f0d876c9e6949beab733 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Wed, 23 Sep 2026 14:57:07 +0200 Subject: [PATCH 17/24] Lots of fixes by Anja Punctuation, Jargon, Capitalization, repeated words and most importantly sphinx warning errors Co-authored-by: Anja Co-authored-by: Michele Mesiti --- content/conclusions.md | 68 ++++++++++++++++++-------------------- content/motivation.md | 4 +-- content/quick-reference.md | 4 +-- content/remotely.md | 10 +++--- 4 files changed, 41 insertions(+), 45 deletions(-) diff --git a/content/conclusions.md b/content/conclusions.md index 7e24aff8..a214f35e 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -26,19 +26,19 @@ after attending this lesson. ## The basics: what knowledge do you need? Learn one test framework well enough for basics: -- Explore and use the good tools that exist out there +- Explore and use the good tools that exist out there. - An incomplete list of {term}`testing framework`s can be found in the [Quick Reference](quick-reference) (see [Unit test frameworks](./quick-reference.md#unit-test-frameworks)). ### Why use a testing framework? -Automated testing typically requires to do a number of repetitive tasks -and to solve some tricky problems. +Automated testing typically involves a number of repetitive tasks +and tricky problem solving. Fortunately for us, someone has already found a solution for most of these -and created {term}`testing framework`s that we can use +and created {term}`testing framework`s that we can use. Note: not all frameworks solve all problems (also because sometimes the underlying language does not have the necessary features). @@ -48,52 +48,50 @@ Note: not all frameworks solve all problems * - **Typical problem** - **Solution** - **Examples** -* - Report failures/successes - (to humans or other machines) - consistently +* - Report failures/successes consistently + (to humans or other machines) - Automated collection and output, (e.g., Junit XML format or [TAP](https://en.wikipedia.org/wiki/Test_Anything_Protocol)) - - Fundamental feature - all testing frameworks - have it. -* - Remember to run all the test you write + - Fundamental feature that + all testing frameworks have +* - Remember to run all the tests you write in the main test script/program - Automatic discovery, automatic {term}`registration` when declaring test functions - - [Pytest](https://docs.pytest.org/en/stable/) + - [pytest](https://docs.pytest.org/en/stable/) * - Run only some tests (to save time) - Test filtering via patterns - - + - [pytest](https://docs.pytest.org/en/stable/how-to/usage.html#specifying-which-tests-to-run) * - Provide useful information on why a test has failed - "Smart" assertions/macros - - [in pytest](https://docs.pytest.org/en/stable/how-to/assert.html#assert), - [in GoogleTest](https://google.github.io/googletest/primer.html#assertions) + - [pytest](https://docs.pytest.org/en/stable/how-to/assert.html#assert), + [GoogleTest](https://google.github.io/googletest/primer.html#assertions) * - Run same test for many known input/output combinations - Parametric tests - - [In Pytest](https://docs.pytest.org/en/stable/how-to/parametrize.html#pytest-mark-parametrize-parametrizing-test-functions), - [in Julia](https://docs.julialang.org/en/v1/stdlib/Test/#Working-with-Test-Sets) (see `testset for`) + - [pytest](https://docs.pytest.org/en/stable/how-to/parametrize.html#pytest-mark-parametrize-parametrizing-test-functions), + [Julia](https://docs.julialang.org/en/v1/stdlib/Test/#Working-with-Test-Sets) (see `testset for`) * - Check that a property holds for a class of inputs and outputs - Automatically generate test cases based on a strategy (property testing) - [hypothesis](https://hypothesis.readthedocs.io/en/latest/tutorial/introduction.html) - (python) + (Python) * - Debugging on failure - Start debugger on test failure - `pytest --pdb` * - Floating point equalities with tolerance - Macros/classes - - `≈` (julia), `pytest.approx` + - `≈` (Julia), `pytest.approx` * - Set up and tear down complex test cases - {term}`Fixture`s - - [In pytest](https://docs.pytest.org/en/stable/explanation/fixtures.html), - [in GoogleTest](https://google.github.io/googletest/primer.html#same-data-multiple-tests) + - [pytest](https://docs.pytest.org/en/stable/explanation/fixtures.html), + [GoogleTest](https://google.github.io/googletest/primer.html#same-data-multiple-tests) * - Estimate how much of your code is *run* in the test suite - Automatic {term}`coverage` @@ -103,7 +101,8 @@ Note: not all frameworks solve all problems (for C/C++/Fortran) * - Do code examples in documentation work as expected? - - doctests + - Documentation tests + (doctests) - [Python](https://docs.python.org/3/library/doctest.html), [Julia](https://documenter.juliadocs.org/stable/man/doctests/), [R](https://cran.r-project.org/web/packages/doctest/vignettes/doctest.html) @@ -111,8 +110,8 @@ Note: not all frameworks solve all problems with different versions of the dependencies? - Test in different environments - - [Tox](https://tox.wiki/en), [nox](https://nox.thea.codes/en/stable/index.html) - (python) + - [tox](https://tox.wiki/en), [Nox](https://nox.thea.codes/en/stable/index.html) + (Python) ``` @@ -128,11 +127,11 @@ what are some of the easy starting points? **If you have got nothing yet**: 1. Start with an end-to-end test. - Typically easy to add, from a typical "manual" use case. + Typically easy to add, from a "manual" use case. This should match (or serve as) an **example in the code documentation** anyway. - Describe in words how *you* check whether the code still works. - Translate the words into a script. - - Run the script as often as reasonable + - Run the script as often as reasonable. 2. Do you have some single functions that are easy to test, but hard to verify just by looking at them? Add unit tests. @@ -142,20 +141,19 @@ what are some of the easy starting points? if you break things. It's actually pretty freeing. **If you need to start modifying some existing code:** -1. Add a {term}`characterization test` for the part of the code you need to change -2. Add tests for any functionality you intend to add +1. Add a {term}`characterization test` for the part of the code you need to change. +2. Add tests for any functionality you intend to add. 3. Consider adding some end-to-end tests for the use case you have in mind. ## Going more in-depth -- Strike a healthy balance between unit tests and integration tests +- Strike a healthy balance between unit tests and integration tests. - As the code gets larger and the chance of undetected bugs - increases, tests should increase -- When adding new functionality, also add tests -- When you discover and fix a bug, also commit a test against this bug + increases, tests should increase. +- When adding new functionality, also add tests. +- When you discover and fix a bug, also commit a test against this bug. - Use {term}`code coverage` analysis to identify untested or unused code. - Remember [Goodhart's Law](https://en.wikipedia.org/wiki/Goodhart%27s_law) -- If you make your code easier to test, it becomes more modular (and vice versa - see the [modular code development lesson](https://coderefinery.github.io/modular-type-along/)) - + Remember [Goodhart's Law](https://en.wikipedia.org/wiki/Goodhart%27s_law). +- If you make your code easier to test, it becomes more modular (and vice versa - see the [modular code development lesson](https://coderefinery.github.io/modular-type-along/)). - **Learning how to test well make the rest of your code better, too.** diff --git a/content/motivation.md b/content/motivation.md index b13ae3e6..e1b9115f 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -92,7 +92,7 @@ With testing, simulations and analysis using software *can* be held to the same Very few people are proud of the code they write the first time they write it. -Typically, they'd like to improve it. +Often, they'd like to improve it. But code without automated tests cannot be improved as easily as code with automated tests. @@ -109,7 +109,7 @@ The [Modular code development lesson](https://coderefinery.github.io/modular-typ There are many forms of testing. -One can write test programs and run them (a form of {term}`end-to-end test`): +One can write test programs and run them (a form of {term}`end-to-end test`ing): ```console $ python3 run-test.py diff --git a/content/quick-reference.md b/content/quick-reference.md index d663d6fb..f6b3b2f4 100644 --- a/content/quick-reference.md +++ b/content/quick-reference.md @@ -25,7 +25,6 @@ End-to-end test and checking that the output is the expected one. See also {term}`Regression test`) - Regression test Check that results and behaviour are what they are supposed to be. @@ -99,13 +98,12 @@ Fixture has run. Property testing - test that a property of the code holds + Test that a property of the code holds for a whole class of inputs. Typically done by automatically generating test cases according to a strategy. Tends to very time-consuming compared to unit testing. - Test Registration The act of marking a test for execution in a main testing program. diff --git a/content/remotely.md b/content/remotely.md index eed6ea75..f0814c64 100644 --- a/content/remotely.md +++ b/content/remotely.md @@ -602,18 +602,18 @@ Finally, we discuss together about our experiences with this exercise. * - Composable actions - [GitHub Marketplace](https://github.com/marketplace?type=actions), very mature ecosystem - - [components](https://gitlab.com/explore/catalog) + - [Components](https://gitlab.com/explore/catalog) * - Defaults - - very minimal + - Very minimal (checkout requires an action) - - typical use cases + - Typical use cases are "baked in" * - Examples and templates - Template workflows - Some [Pipeline examples](https://docs.gitlab.com/ci/examples/) * - Execution - - [GitHub-hosted Runners](https://docs.github.com/en/actions/concepts/runners/github-hosted-runners) - - [GitLab-hosted Runners](https://docs.gitlab.com/ci/runners/#gitlab-hosted-runners) + - [GitHub-hosted runners](https://docs.github.com/en/actions/concepts/runners/github-hosted-runners) + - [GitLab-hosted runners](https://docs.gitlab.com/ci/runners/#gitlab-hosted-runners) * - Self-hosted execution - [Self-hosted runners](https://docs.github.com/en/actions/how-tos/manage-runners/self-hosted-runners) From c70cdf94cb097a9785fa046b1595c6df873e38e4 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Wed, 23 Sep 2026 17:15:19 +0200 Subject: [PATCH 18/24] further sphinx warning fixes --- content/conclusions.md | 2 +- content/quick-reference.md | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/content/conclusions.md b/content/conclusions.md index a214f35e..d09bcb41 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -29,7 +29,7 @@ Learn one test framework well enough for basics: - Explore and use the good tools that exist out there. - An incomplete list of {term}`testing framework`s can be found in the [Quick Reference](quick-reference) - (see [Unit test frameworks](./quick-reference.md#unit-test-frameworks)). + (see [Unit test frameworks](#unit-test-frameworks)). ### Why use a testing framework? diff --git a/content/quick-reference.md b/content/quick-reference.md index f6b3b2f4..cbf8179c 100644 --- a/content/quick-reference.md +++ b/content/quick-reference.md @@ -104,6 +104,7 @@ Property testing test cases according to a strategy. Tends to very time-consuming compared to unit testing. + Test Registration The act of marking a test for execution in a main testing program. @@ -118,6 +119,7 @@ Test Registration ## Available tools +(unit-test-frameworks)= ### Unit test frameworks A **test framework** makes it easy to run tests across large amounts From 906dea22c837eb062b4ced0f2cbc2763ab917218 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Wed, 23 Sep 2026 19:19:49 +0200 Subject: [PATCH 19/24] (minor) terser / better wording also, made title of forge feature comparison table more sober. --- content/remotely.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/content/remotely.md b/content/remotely.md index f0814c64..0950bc7a 100644 --- a/content/remotely.md +++ b/content/remotely.md @@ -593,7 +593,7 @@ Finally, we discuss together about our experiences with this exercise. **There is more tooling available** to discover, for example: -```{list-table} A jungle of tools +```{list-table} Features and implementations on GitHub Actions and GitLab CI/CD * - **Feature** - **GitHub** **Actions** @@ -624,13 +624,13 @@ Finally, we discuss together about our experiences with this exercise. Both GitHub Actions (on *github.com*) and GitLab CI/CD (on *gitlab.com*) offer runners on Linux, Windows and MacOS machines. -If you need more control, you can choose to **self-host**: +If you need more control, you can **self-host**: - CI/CD and Actions can also be made available for self-hosted GitLab servers or GitHub Enterprise (i.e., outside *gitlab.com* or *github.com* domains). Your data does not have to go to a cloud! - If you need more control on the running environment, - you can also connect the git server + you can connect the git server to another host (including, e.g., a HPC system) and run workflows/pipelines there, From 77f7b50f7fa198cfa74194b26f132ddcc74ef3a4 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Wed, 23 Sep 2026 19:26:03 +0200 Subject: [PATCH 20/24] conclusions: Improve wording on knowledge needed Co-Authored-by: Anja Virkkunen --- content/conclusions.md | 7 +++---- 1 file changed, 3 insertions(+), 4 deletions(-) diff --git a/content/conclusions.md b/content/conclusions.md index d09bcb41..5e0d86e2 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -25,11 +25,10 @@ after attending this lesson. ## The basics: what knowledge do you need? -Learn one test framework well enough for basics: +Learn one {term}`testing framework` well enough for basics: - Explore and use the good tools that exist out there. -- An incomplete list of {term}`testing framework`s - can be found in the [Quick Reference](quick-reference) - (see [Unit test frameworks](#unit-test-frameworks)). +- An incomplete list of testing frameworks + can be found [here](#unit-test-frameworks). ### Why use a testing framework? From 084c7a1ec2cec5b318692a1f537ed79018d2f7fc Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Wed, 23 Sep 2026 19:32:27 +0200 Subject: [PATCH 21/24] conclusions: test framework feat table col fixes slimmer columns for readability --- content/conclusions.md | 17 +++++++++++------ 1 file changed, 11 insertions(+), 6 deletions(-) diff --git a/content/conclusions.md b/content/conclusions.md index 5e0d86e2..ae72bcac 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -47,13 +47,15 @@ Note: not all frameworks solve all problems * - **Typical problem** - **Solution** - **Examples** -* - Report failures/successes consistently +* - Report failures/successes + consistently (to humans or other machines) - Automated collection and output, (e.g., Junit XML format or [TAP](https://en.wikipedia.org/wiki/Test_Anything_Protocol)) - - Fundamental feature that + - Fundamental feature that all testing frameworks have -* - Remember to run all the tests you write +* - Remember to run + all the tests you write in the main test script/program - Automatic discovery, automatic {term}`registration` @@ -84,10 +86,12 @@ Note: not all frameworks solve all problems * - Debugging on failure - Start debugger on test failure - `pytest --pdb` -* - Floating point equalities with tolerance +* - Floating point equalities + with tolerance - Macros/classes - `≈` (Julia), `pytest.approx` -* - Set up and tear down complex test cases +* - Set up and tear down + of complex test cases - {term}`Fixture`s - [pytest](https://docs.pytest.org/en/stable/explanation/fixtures.html), [GoogleTest](https://google.github.io/googletest/primer.html#same-data-multiple-tests) @@ -98,7 +102,8 @@ Note: not all frameworks solve all problems - [Pytest-cov](https://pytest-cov.readthedocs.io/en/latest/), [gcov/lcov](https://wiki.cs.jmu.edu/reference/gcov/) (for C/C++/Fortran) -* - Do code examples in documentation +* - Do code examples + in documentation work as expected? - Documentation tests (doctests) From a187734f5cd3806902d4cd554ab9e84f2a1bbe67 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Thu, 24 Sep 2026 19:42:25 +0200 Subject: [PATCH 22/24] Mention the test pyramid Adding a picture here might be unnecessary since there's so many that one can find online. --- content/conclusions.md | 21 ++++++++++++++++----- content/quick-reference.md | 11 +++++++++++ 2 files changed, 27 insertions(+), 5 deletions(-) diff --git a/content/conclusions.md b/content/conclusions.md index ae72bcac..89193a41 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -151,13 +151,24 @@ what are some of the easy starting points? ## Going more in-depth -- Strike a healthy balance between unit tests and integration tests. -- As the code gets larger and the chance of undetected bugs - increases, tests should increase. -- When adding new functionality, also add tests. -- When you discover and fix a bug, also commit a test against this bug. +**With time**: +- The code gets larger, + the chance of undetected bugs increases, + tests should increase. +- Bugs will be found. When you find them, add tests against those. + + +**How to improve your code:** - Use {term}`code coverage` analysis to identify untested or unused code. Remember [Goodhart's Law](https://en.wikipedia.org/wiki/Goodhart%27s_law). +- Strike a healthy balance between different kinds of tests: + - Fast tests give you information quicker but can be shallow; + - Thorough tests can catch more bugs but take longer to run + and can be brittle. + + The {term}`test pyramid` + is a recommended strategy to balance between test types. + - If you make your code easier to test, it becomes more modular (and vice versa - see the [modular code development lesson](https://coderefinery.github.io/modular-type-along/)). - **Learning how to test well make the rest of your code better, too.** diff --git a/content/quick-reference.md b/content/quick-reference.md index cbf8179c..b1646602 100644 --- a/content/quick-reference.md +++ b/content/quick-reference.md @@ -115,6 +115,17 @@ Test Registration in the "main" script/function, with the risk of forgetting it. +Test Pyramid + The general approach for balancing test types + in a test suite. + The slower a test type is, + the fewere tests of that type + should be in the test suite. + + Search engines can show many representations of this. + + + ::::: ## Available tools From ce9039ab2ac463280d7caaefb556c3ceb754efec Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Thu, 24 Sep 2026 19:52:18 +0200 Subject: [PATCH 23/24] Split 'Exercise' in local testing in 2 sections --- content/locally.md | 10 +++++++++- 1 file changed, 9 insertions(+), 1 deletion(-) diff --git a/content/locally.md b/content/locally.md index 0b598d50..24d9bf37 100644 --- a/content/locally.md +++ b/content/locally.md @@ -5,7 +5,7 @@ ``` -## Exercise +## Setting up your first automated test In this exercise we will make a simple function and use @@ -267,6 +267,14 @@ whether our test detects the change: ``````` ````````` +## Numerical Tolerances + +Some times the testing logic needs to be slightly more complicated. +In scientific computing +many functions return floating point numbers: +how do we test them? + + `````````{challenge} (optional) Local-2: Create a test that considers numerical tolerance (10 min) Let's see an example where the test has to be more clever in order to avoid false negative. From dfb8fb7ff1ec2aa533ed13b12ca91734048119d1 Mon Sep 17 00:00:00 2001 From: Michele Mesiti Date: Thu, 24 Sep 2026 19:52:57 +0200 Subject: [PATCH 24/24] local testing: false negative -> false positive false positive is a bug is flagged when not present false negative is when a bug is there but it's not flagged. Notice that tests can only be used to prove incorrectness, not to prove correctness, so the condition we are testing for is the presence of bugs, not its absence. --- content/locally.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/content/locally.md b/content/locally.md index 24d9bf37..c55ba391 100644 --- a/content/locally.md +++ b/content/locally.md @@ -277,7 +277,7 @@ how do we test them? `````````{challenge} (optional) Local-2: Create a test that considers numerical tolerance (10 min) Let's see an example where the test has to be more clever in order to -avoid false negative. +avoid false positive. In the above exercise we have compared integers. In this optional exercise we want to learn how to compare floating point numbers since they are more tricky