diff --git a/content/conclusions.md b/content/conclusions.md index d2872a7..89193a4 100644 --- a/content/conclusions.md +++ b/content/conclusions.md @@ -1,39 +1,174 @@ # Conclusions and recommendations -## The basics +```{objectives} +- What is a realistic approach to automated test? +- How can I start? +``` -- Learn one test framework well enough for basics - - Explore and use the good tools that exist out there - - An incomplete list of testing frameworks can be found in the [Quick Reference](quick-reference) -- Start with some basics - - Some simple thing that test all parts -- Automate tests - - Faster feedback and reduce the number of surprises +--- +## Discussion: What's easy and hard to test? -## Going more in-depth +```{discussion} Discussion: Testing in practice + +Use the collaborative notes to answer these questions: + +1. Give examples of things (from your work) that are easy to test. +2. Give examples of things (from your work) that are hard to test. +``` + +--- + +Considering automated tests when writing code +is a major mental shift, that you will hopefully embrace +after attending this lesson. + +## The basics: what knowledge do you need? + +Learn one {term}`testing framework` well enough for basics: +- Explore and use the good tools that exist out there. +- An incomplete list of testing frameworks + can be found [here](#unit-test-frameworks). + +### Why use a testing framework? + +Automated testing typically involves a number of repetitive tasks +and tricky problem solving. + +Fortunately for us, +someone has already found a solution for most of these +and created {term}`testing framework`s that we can use. + +Note: not all frameworks solve all problems +(also because sometimes the underlying language does not have the necessary features). + +```{list-table} Why use a testing framework? +:widths: 40 40 20 +* - **Typical problem** + - **Solution** + - **Examples** +* - Report failures/successes + consistently + (to humans or other machines) + - Automated collection and output, + (e.g., Junit XML format or [TAP](https://en.wikipedia.org/wiki/Test_Anything_Protocol)) + - Fundamental feature that + all testing frameworks have +* - Remember to run + all the tests you write + in the main test script/program + - Automatic discovery, + automatic {term}`registration` + when declaring test functions + - [pytest](https://docs.pytest.org/en/stable/) +* - Run only some tests + (to save time) + - Test filtering + via patterns + - [pytest](https://docs.pytest.org/en/stable/how-to/usage.html#specifying-which-tests-to-run) +* - Provide useful information + on why a test has failed + - "Smart" assertions/macros + - [pytest](https://docs.pytest.org/en/stable/how-to/assert.html#assert), + [GoogleTest](https://google.github.io/googletest/primer.html#assertions) +* - Run same test for many + known input/output combinations + - Parametric tests + - [pytest](https://docs.pytest.org/en/stable/how-to/parametrize.html#pytest-mark-parametrize-parametrizing-test-functions), + [Julia](https://docs.julialang.org/en/v1/stdlib/Test/#Working-with-Test-Sets) (see `testset for`) +* - Check that a property holds + for a class of inputs and outputs + - Automatically generate test cases + based on a strategy + (property testing) + - [hypothesis](https://hypothesis.readthedocs.io/en/latest/tutorial/introduction.html) + (Python) +* - Debugging on failure + - Start debugger on test failure + - `pytest --pdb` +* - Floating point equalities + with tolerance + - Macros/classes + - `≈` (Julia), `pytest.approx` +* - Set up and tear down + of complex test cases + - {term}`Fixture`s + - [pytest](https://docs.pytest.org/en/stable/explanation/fixtures.html), + [GoogleTest](https://google.github.io/googletest/primer.html#same-data-multiple-tests) +* - Estimate how much of your code + is *run* in the test suite + - Automatic {term}`coverage` + measurement + - [Pytest-cov](https://pytest-cov.readthedocs.io/en/latest/), + [gcov/lcov](https://wiki.cs.jmu.edu/reference/gcov/) + (for C/C++/Fortran) +* - Do code examples + in documentation + work as expected? + - Documentation tests + (doctests) + - [Python](https://docs.python.org/3/library/doctest.html), + [Julia](https://documenter.juliadocs.org/stable/man/doctests/), + [R](https://cran.r-project.org/web/packages/doctest/vignettes/doctest.html) +* - Will my code work + with different versions + of the dependencies? + - Test in different environments + - [tox](https://tox.wiki/en), [Nox](https://nox.thea.codes/en/stable/index.html) + (Python) +``` -- Strike a healthy balance between unit tests and integration tests -- As the code gets larger and the chance of undetected bugs - increases, tests should increase -- When adding new functionality, also add tests -- When you discover and fix a bug, also commit a test against this bug -- Use code coverage analysis to identify untested or unused code -- If you make your code easier to test, it becomes more modular +## Don't over-test -## Ways to get started +- Not every code needs perfect {term}`test coverage`. +- A simple script or notebook probably does not need an automated test. + +## Pick the low-hanging fruits first You probably won't do everything perfectly when you start off... But what are some of the easy starting points? -- Do you have some single functions that are easy to test, but hard to - verify just by looking at them? Add unit tests. +**If you have got nothing yet**: +1. Start with an end-to-end test. + Typically easy to add, from a "manual" use case. + This should match (or serve as) an **example in the code documentation** anyway. + - Describe in words how *you* check whether the code still works. + - Translate the words into a script. + - Run the script as often as reasonable. -- Do you have data analysis or simulation of some sort? Make an - end-to-end test with sample data, or sample parameters. This is - useful as an example anyway. +2. Do you have some single functions that are easy to test, but hard to + verify just by looking at them? Add unit tests. -- A local testing framework + GitHub actions is very easy! And works - well in the background - you do whatever you want and get an email +3. A local testing framework + GitHub actions/Gitlab CI-CD is very easy! + And works well in the background - you do whatever you want and get an email if you break things. It's actually pretty freeing. + +**If you need to start modifying some existing code:** +1. Add a {term}`characterization test` for the part of the code you need to change. +2. Add tests for any functionality you intend to add. +3. Consider adding some end-to-end tests for the use case you have in mind. + +## Going more in-depth + +**With time**: +- The code gets larger, + the chance of undetected bugs increases, + tests should increase. +- Bugs will be found. When you find them, add tests against those. + + +**How to improve your code:** +- Use {term}`code coverage` analysis to identify untested or unused code. + Remember [Goodhart's Law](https://en.wikipedia.org/wiki/Goodhart%27s_law). +- Strike a healthy balance between different kinds of tests: + - Fast tests give you information quicker but can be shallow; + - Thorough tests can catch more bugs but take longer to run + and can be brittle. + + The {term}`test pyramid` + is a recommended strategy to balance between test types. + +- If you make your code easier to test, it becomes more modular (and vice versa - see the [modular code development lesson](https://coderefinery.github.io/modular-type-along/)). +- **Learning how to test well make the rest of your code better, too.** + diff --git a/content/locally.md b/content/locally.md index 0b598d5..c55ba39 100644 --- a/content/locally.md +++ b/content/locally.md @@ -5,7 +5,7 @@ ``` -## Exercise +## Setting up your first automated test In this exercise we will make a simple function and use @@ -267,9 +267,17 @@ whether our test detects the change: ``````` ````````` +## Numerical Tolerances + +Some times the testing logic needs to be slightly more complicated. +In scientific computing +many functions return floating point numbers: +how do we test them? + + `````````{challenge} (optional) Local-2: Create a test that considers numerical tolerance (10 min) Let's see an example where the test has to be more clever in order to -avoid false negative. +avoid false positive. In the above exercise we have compared integers. In this optional exercise we want to learn how to compare floating point numbers since they are more tricky diff --git a/content/motivation.md b/content/motivation.md index 7468fe5..e1b9115 100644 --- a/content/motivation.md +++ b/content/motivation.md @@ -5,8 +5,7 @@ - Understand various benefits of testing ``` -Most scientists nowadays depend on software for research. - +Most scientists nowadays depend on software for research. What can go wrong when research software has bugs? Look no further: - [A Scientist's Nightmare: Software Problem Leads to Five Retractions](https://science.sciencemag.org/content/314/5807/1856.summary) @@ -16,7 +15,7 @@ How can we avoid problems like these? ## What are typical problems that *automated* tests can address? -Have you ever had some of these problems? +Have you ever had any of these problems? - You change B and C, and suddenly A doesn't work anymore. Time wasted trying to figure out what changed. @@ -40,23 +39,89 @@ places it's useful for research code, and how easy it can be. ## Untested software can be compared to uncalibrated measurement devices -*"Before relying on a new experimental device, an experimental scientist always +```{epigraph} +*Before relying on a new experimental device, an experimental scientist always establishes its accuracy. A new detector is calibrated when the scientist observes its responses to known input signals. The results of this -calibration are compared against the expected response."* +calibration are compared against the expected response.* -> [From [Testing and Continuous Integration with Python](https://carpentries-incubator.github.io/python-testing/), created by K. Huff] +-- From [Testing and Continuous Integration with Python](https://carpentries-incubator.github.io/python-testing/), created by K. Huff +``` With testing, simulations and analysis using software *can* be held to the same standards as experimental measurement devices! --- +## What can tests help you do? + +```{list-table} Problems, Solution and who is affected? +:widths: 40 30 30 +* - **Problem** + - **Solution** + - **Who is affected?** +* - Breaking old functionality + when adding new features + - {term}`End-to-End test`s + - Developers +* - Verify installation + - {term}`Smoke test`s + - Users +* - Make small incremental changes, + (e.g., improving readability, names) + - {term}`Unit test`s + - Developers +* - Make architectural changes, + (e.g., shifting code between + classes, modules and functions) + - {term}`Integration test`s + {term}`End-to-End test`s + - Developers +* - Change things with confidence + that nothing is breaking + - All tests + - Developers +* - Documentation out of date + including code examples + - Executable notebooks + and [nbval](https://github.com/computationalmodelling/nbval), + {term}`End-to-End test`s + - Users + +``` + +Very few people are proud of the code they write +the first time they write it. + +Often, they'd like to improve it. + +But code without automated tests cannot be improved as easily +as code with automated tests. + +Moreover, **code that is easy to test is probably easier to maintain**, +since it needs to be more modular and have better separation of concerns. + +The [Modular code development lesson](https://coderefinery.github.io/modular-type-along/) + demonstrates this. + +--- + ## Testing in a nutshell -In the most basic form of software tests, -expected results are compared with observed results -in order to establish accuracy. Why are we not comparing directly all -digits with the expected result? +There are many forms of testing. + +One can write test programs and run them (a form of {term}`end-to-end test`ing): +```console +$ python3 run-test.py + +running: sample_data/set1.csv --output=tests/set1.txt +CORRECT +``` + +In the most basic form of a software test, +the observed result is compared with expected result (an "*oracle*") +in order to establish correctness. +Here are some examples of this testing pattern +in different programming languages (in this case, {term}`unit test`s): ````{tabs} ```{group-tab} Python @@ -95,61 +160,12 @@ digits with the expected result? ``` ```` -Or you can test whole programs: -```console -$ python3 run-test.py - -running: sample_data/set1.csv --output=tests/set1.txt -CORRECT -``` - ---- - -## What can tests help you do? - -```{list-table} Problems, Solutions and who is affected? -:widths: 40 30 30 -* - Problem - - Solution - - Who is affected? -* - Breaking old functionality - when adding new features - - End-to-End tests - - Developers -* - Verify installation - - Smoke tests - - Users -* - Showing up-to-date example - - End-to-End tests - - Users -* - Improve readability and names - - Unit tests - - Developers -* - Refactor and restructure - - All tests - - Developers -* - Documentation out of date - - Executable notebooks - and [nbval](https://github.com/computationalmodelling/nbval), - End-to-End tests - - Users - -``` - -**Help other developers modify it** -- Change things with confidence that nothing is breaking. -- Warning if documentation/examples go out of date. - -**Manage complexity** -- If code is easy to test, it's probably easier to maintain. -- The next lesson [Modular code development](https://coderefinery.github.io/modular-type-along/) - demonstrates this. --- -## Discussion: When is it OK not to add tests? +## Discussion: When is it OK not to add automated tests? -```{discussion} Discussion: When is it OK not to add tests? +::::{discussion} Discussion: When is it OK not to add automated tests? Vote in the notes and we'll discuss soon. **It is always a balance: there is no "always"/"never"**. @@ -159,59 +175,29 @@ Vote in the notes and we'll discuss soon. **It is always a balance: there is no 3. A simple short, "obviously correct" shell script? 4. Can you give other examples? -``` +:::{solution} +The role of automated tests is to save time when making changes to code. ---- +1. In this case you just "test manually" the notebook by running it. Automated tests might not save you time. + But if some non trivial functions are added, you might want to have automated {term}`unit test`s for these separately. -## Discussion: What's easy and hard to test? +2. Writing automated tests for "throwaway code" can be a waste of time. But if you get back to it, then you should think about writing automated tests for it. -```{discussion} Discussion: Testing in practice +3. "Manual test" can be sufficient. In case of changes, checking the script with a {term}`linter` like [Shellcheck](https://www.shellcheck.net/) might still be useful! -Use the collaborative notes to answer these questions: +::: -1. Give examples of things (from your work) that are easy to test. -2. Give examples of things (from your work) that are hard to test. -``` - ---- - -## Testing vocabulary - -* Test functions and methods one at a time - **Unit tests** - -* Test how parts work together - **Integration tests** - -* Test the whole thing running, checking the output - **End-to-end tests** - * For example, running on sample data. - -* Test that the whole thing runs in the simplest scenario possible - **Smoke Test** - * if this fails, no point in testing other things, usually. - -* Check results are the same as before - **Regression tests** - * Other names for the same thing: **Acceptance Tests**, **Golden-Master Tests**, **Characterization Tests** - -* Write test first (the output), then write code to make test pass - - **Test-driven development** - -* GitHub or GitLab runs tests automatically - **Continuous - integration** - -* Report that tells you which lines were/were not run by tests - - **Code coverage** - -* Framework that runs test for you - **Testing framework** - * See [Quick Reference](./quick-reference) for some examples. +:::: --- ## What should you do? -* Not every code needs perfect test coverage. * If code is interactive-only (Jupyter Notebook), it's usually hard to test. - * But also hard to run: the next lesson will discuss! + * But also hard to run: the [next lesson](https://coderefinery.github.io/modular-type-along/) will discuss! * At least end-to-end is often easy to add. @@ -223,21 +209,8 @@ Use the collaborative notes to answer these questions: * It's easy to have Gitlab/Github run the tests. * It's nice to push without thinking, and the system tells you when - it's broke. + it's broken. * **Learning how to test well make the rest of your code better, too.** ---- - -## Where to start - -- A simple script or notebook probably does not need an automated test. - -**If you have nothing yet** -- Start with an end-to-end test. -- Describe in words how *you* check whether the code still works. -- Translate the words into a script. -- Run the script automatically on every code change. -**If you want to start with unit-testing** -- You want to rewrite a function? Start adding a unit test right there first. diff --git a/content/quick-reference.md b/content/quick-reference.md index 3cff9f4..b164660 100644 --- a/content/quick-reference.md +++ b/content/quick-reference.md @@ -1,7 +1,136 @@ # Quick Reference +## Glossary +:::::{glossary} + +Unit test + A test that covers a single functions or method + (a "unit") + +Integration test + A test that checks if "units" work together as intended + +Smoke Test + Check that the whole application or script runs without errors + in the simplest scenario possible. + If this fails, no point in testing other things, usually + + +Regression + A loss of functionality, typically due to a bug + +End-to-end test + Test the whole thing running, checking the output + (For example, running on sample data + and checking that the output is the expected one. + See also {term}`Regression test`) + +Regression test + + Check that results and behaviour are what they are supposed to be. + Typically written once a {term}`regression` is detected. + Other names for the same kind of test: + - Characterization Tests + - Acceptance Tests + - Golden-Master Tests + +Characterization Test + + Automated test that is written on existing (legacy) code, + (assuming that version of the code is correct) + and added to a test suite + to make further work/changes easier + +Test-first development + + The practice of writing automated tests + before writing the code that makes the tests pass + +TDD + Acronym for {term}`Test-driven development` + +Test-driven development + A special case of Test-First development + where the workflow is: + + - Make a list of specifications + + - Then, for each specification: + - Write a test, run it and verify that the test fails + - Write the minimum amount of code to make the test pass + - Refactor and improve the code + + +Continuous integration + + The practice of merging in the main branch frequently + without having long-lived branches. + This typically requires automating part of the workflow, + especially testing, + and GitHub/GitLab et similia have support for that + via Actions/CI-CD respectively. + +Code coverage + + Metric representing the fraction of code base + executed during the test suite. + Note: this is only an **upper bound** + to the fraction of code base + that is anyhow tested. + And it is perfectly possible to write tests that are completely useless + but increase code coverage. + +Testing framework + + Framework that runs tests for you. + See the following for some examples. + +Linter + + A program that can check your code + for typical mistakes + or for risky practices, and reports them to you. + +Fixture + A resource that needs to be set up before a test case can run + and needs to be torn down after the test case + (or a whole test suite) + has run. + +Property testing + Test that a property of the code holds + for a whole class of inputs. + Typically done by automatically generating + test cases according to a strategy. + Tends to very time-consuming + compared to unit testing. + +Test Registration + The act of marking a test for execution + in a main testing program. + Test frameworks allow to do this + automatically + at test definition + so that it does not have to be manually invoked + in the "main" script/function, + with the risk of forgetting it. + +Test Pyramid + The general approach for balancing test types + in a test suite. + The slower a test type is, + the fewere tests of that type + should be in the test suite. + + Search engines can show many representations of this. + + + +::::: + ## Available tools +(unit-test-frameworks)= ### Unit test frameworks A **test framework** makes it easy to run tests across large amounts @@ -281,10 +410,8 @@ You can then compile using this script: Each of these are web services to handle testing, free for open source projects. -- [GitHub Actions](https://github.com/features/actions) (we will - demonstrate this in the next episode) -- [GitLab CI](https://about.gitlab.com/features/gitlab-ci-cd/) - (we will demonstrate this in the next episode) +- [GitHub Actions](https://github.com/features/actions) - see episode [Automated Testing Remotely](./remotely) +- [GitLab CI](https://about.gitlab.com/features/gitlab-ci-cd/) - see episode [Automated Testing Remotely](./remotely) - [Azure Pipelines](https://azure.microsoft.com/en-us/services/devops/pipelines/) - [Coveralls](https://coveralls.io) - [Codecov](https://codecov.io) diff --git a/content/remotely.md b/content/remotely.md index f75abf4..0950bc7 100644 --- a/content/remotely.md +++ b/content/remotely.md @@ -27,7 +27,7 @@ In this exercise, we will: - **C.** Find a bug in our repository and open an issue to report it - **D.** Fix the bug on a bugfix branch and open a pull request (GitHub)/ merge request (GitLab) - **E.** Merge the pull/merge request and see how the issue is automatically closed. -- **F.** Create a test to increase the code coverage of our tests. +- **F.** Create a test to increase the {term}`code coverage` of our tests. ``` ## Prerequisites @@ -579,15 +579,64 @@ Finally, we discuss together about our experiences with this exercise. ## Where to go from here -- This example was using Python but you can achieve the same automation for R or Fortran or C/C++ or other languages -- This workflow is very useful for collaborators who work on the same code and it works both for +**These techniques:** +- **apply to all programming languages**: + This example was using Python but you can achieve the same automation for any other languages +- **are recommended and useful for collaborative software development**: + automatically running the test suite remotely + as presented here + is very useful for collaborators who work on the same code and it works both for [centralized](https://coderefinery.github.io/git-collaborative/02-centralized/) and [forking](https://coderefinery.github.io/git-collaborative/03-distributed/) workflows - have a look at this [alternative exercise](./full-cycle-ci) to see how that works. -- GitHub Actions has a [Marketplace](https://github.com/marketplace?type=actions) which offer wide range of automatic workflows -- On GitLab use [GitLab CI](https://about.gitlab.com/product/continuous-integration/) -- For Windows builds you can also use [Appveyor](https://www.appveyor.com) + +**There is more tooling available** to discover, for example: + + +```{list-table} Features and implementations on GitHub Actions and GitLab CI/CD +* - **Feature** + - **GitHub** + **Actions** + - **GitLab** + **CI/CD** +* - Composable actions + - [GitHub Marketplace](https://github.com/marketplace?type=actions), + very mature ecosystem + - [Components](https://gitlab.com/explore/catalog) +* - Defaults + - Very minimal + (checkout requires an action) + - Typical use cases + are "baked in" +* - Examples and templates + - Template workflows + - Some [Pipeline examples](https://docs.gitlab.com/ci/examples/) +* - Execution + - [GitHub-hosted runners](https://docs.github.com/en/actions/concepts/runners/github-hosted-runners) + - [GitLab-hosted runners](https://docs.gitlab.com/ci/runners/#gitlab-hosted-runners) +* - Self-hosted + execution + - [Self-hosted runners](https://docs.github.com/en/actions/how-tos/manage-runners/self-hosted-runners) + - [GitLab runner](https://docs.gitlab.com/runner/) + (more customizable) +``` +Both GitHub Actions (on *github.com*) and GitLab CI/CD (on *gitlab.com*) offer runners on Linux, Windows and MacOS machines. + + +If you need more control, you can **self-host**: +- CI/CD and Actions can also be made available + for self-hosted GitLab servers or GitHub Enterprise + (i.e., outside *gitlab.com* or *github.com* domains). + Your data does not have to go to a cloud! +- If you need more control on the running environment, + you can connect the git server + to another host + (including, e.g., a HPC system) + and run workflows/pipelines there, + with self-hosted runners + (for [GitHub](https://docs.github.com/en/actions/concepts/runners/self-hosted-runners) + and for [GitLab](https://docs.gitlab.com/runner/)). ```{keypoints} - When fixing bugs or other problems reported in issues, use the issue