Dockerized Setup for the MinHash-based Code Recognition and Investigation Toolkit (MCRIT).
This repository intends to enable you to quickly run a production-ready deployment of MCRIT including its frontend MCRITweb with minimal effort through a pre-configured Docker setup.
The latest commit on this repository will always hold references to the most recent versions of the front- and backend, and stable milestone releases will be marked as such.
Given an installation of docker-compose, running this command in the repository root:
$ docker-compose upshould build the MCRIT server and worker as well as the MCRITweb images, pull images for mongodb and nginx, and then start up containers for everything.
The data produced and stored by the services are found in ./storage:
./storage/mcritwebcontains the sqlite DB for the web application and cached data, such as matching reports in JSON format./storage/mongodbcontains all collections and indices, allowing it to persist across MCRIT web instance rebuilds and updates.
By default, the NGINX included in this setup is only listening for the server_name of localhost, so you will need to configure
./nginx/mcritweb_plain.confor recommendably:./nginx/mcritweb_ssl.conf
based on the specifics of your server.
If you want to run the service over HTTPS
- you will need to adjust the NGINX service of the
docker-compose.ymlto use the./nginx/mcritweb_ssl.confinstead of./nginx/mcritweb_plain.confand - Fill the respective files in
./nginx/sslwith a certificate, private key, and ideally fresh Diffie-Hellman parameters.
If you want to use this Docker setup for development on MCRIT, you will need the code repositories available outside of the containers to trivially reflect your changes.
For this, mcrit and mcritweb should first be cloned into ./repositories, for which you can conveniently use the script clone_repositories.sh.
Afterwards, you can start the setup up in development mode, using:
$ docker-compose -f docker-compose-dev.yml upNote that running in development mode will not start up NGINX, meaning you can reach MCRIT only via ports 5000 (frontend) and 8000 (backend).
For an explanation of the usage of MCRIT itself, please refer to the respective repositories for
- backend: MCRIT (documentation in preparation)
- frontend: MCRITweb (documentation)
MCRIT 1.7.0 keeps function disassembly in its own xcfg collection instead of inline in every
function document, so the queries on the matching path stop reading past a blob they do not need.
Upgrading does not require this - the readers fall back to an inline blob when none has been split
out - but until the migration runs, the instance keeps the old layout and none of the benefit.
$ ./migrate.shThe script runs the migration inside the mcrit-server container and then verifies it. It is
resumable, so an interrupted run costs only the batch in flight, and running it again after
completion does nothing. For reference, 11.6M functions took about 17 minutes.
Three things worth knowing before you start:
- it needs free disk roughly the size of your stored disassembly, and does not return it until
you compact the collection afterwards:
docker compose exec mongodb mongosh mcrit --eval 'db.runCommand({compact: "functions"})'(check your MongoDB version's notes -compactblocks operations on the collection in some versions) - byte-for-byte verification is only possible beforehand. An in-place run leaves no original
blobs to compare against, so verification afterwards establishes coverage and counts rather than
content. To get content-level proof, rehearse into a second database first
(
--mode copy --target <db>then--mode verify --target <db>) or take a dump. - do not downgrade a migrated instance without running
--mode unsplitfirst. Older MCRIT versions do not know about thexcfgcollection and fail quietly rather than loudly.
/status reports inline_xcfg_remaining, so an instance that has not finished migrating says so.
Full details in the upstream migration guide.
MCRIT does not pin SMDA, so rebuilding the images can pull a newer SMDA whose instruction escaping differs. That matters because MinHashes are derived from the escaped representation: when the escaping changes, previously indexed signatures stop being comparable to newly computed ones. They still look valid and still match each other, while identical code submitted afterwards no longer finds them. SMDA 4.4.5 is a worked example - it corrected how segment-qualified memory operands are escaped, which altered about a fifth of the stored MinHashes on a real corpus.
MCRIT 1.7.0 makes this visible. /status reports smda_version and escaper_fingerprint:
$ curl -s http://127.0.0.1:8000/status | python -m json.tool | grep -E "smda_version|escaper"Note the fingerprint after a build. If it changes on a later rebuild, the escaping changed, and the index should be re-minhashed to stay internally consistent - the stored disassembly makes that a local recomputation, with no need to re-submit any samples.
MCRIT was officially released as version 1.0.0 at Botconf 2023 (paper, slides, video)
- 2026-08-25: MCRIT 1.8.1 (packaging and latent-bug release, no migration, no re-index and no database shape change; matching results are unchanged. 1.8.0 is skipped deliberately - it does not declare
packaging, whichMongoDbStorageimports, so a server or worker built from it cannot start; 1.8.1 declares it. Rebuild the mcrit images rather than pulling code - MCRIT no longer shipsrequirements.txt, its dependencies are declared inpyproject.toml, and the Dockerfile installs them in one step accordingly; a build against the old Dockerfile fails on the missing file. The shippedconfig/copies are regenerated from 1.8.0 stock, so merge rather than overwrite if you have edited them locally - the only intentional deviations areSTORAGE_SERVER/QUEUE_SERVERpointing at themongodbservice andBAND_MATCHES_REQUIRED = 1, which is left as it was and now says so in a comment.AUTH_TOKENcan now be supplied through theMCRIT_AUTH_TOKENenvironment variable instead of being edited into the config, and the API token is compared in constant time - note that actually locking the API down in this deployment also needs MCRITweb's matching apitoken set, so the variable alone changes nothing here, and the API is only bound to127.0.0.1regardless. Adopting thetytype checker in CI turned up eleven latent bugs, all fixed: fiveMemoryStoragemethods that raised on every call, a corrupt pichash entry written on import, a server that could not start where gunicorn is absent, and a query-sample deletion that reported success as failure (mcrit#102). Malformed search queries answer HTTP 400 with the parser's message instead of a 500 and a traceback, and unique-block requests reportyara_coversfor real rather than always0), MCRITweb 1.4.8 - 2026-08-24: MCRIT 1.7.1 (bugfix release, no migration, no database shape change and no new settings - the shipped
config/copies are unchanged apart fromVERSION, so a local merge is not needed this time. Search conditions onpichashare served instead of raising:=,!=and substring searches work, while range comparisons such aspichash:<0x99are rejected with a message, because the stored form is variable-width hex on which comparisons are not meaningful (see mcrit#145). The search endpoints now answer an unsupported query with HTTP 400 and that message instead of a 500 and a traceback;McritClienttreats both as no result, so nothing downstream changes. Cross-compare reports no longer emit phantom rows for samples outside the requested set. Sorting a function search by pichash is now refused on the MongoDB backend rather than returning a mis-sorted first page and failing on the second - the MCRITweb function table has no sortable Pichash column, so the UI here is unaffected and only a client calling the API withsort_by=pichashnotices), MCRITweb 1.4.8 - 2026-08-24: MCRIT 1.7.0 (database shape change: disassembly moves out of the function documents into dedicated
xcfg/query_xcfgcollections - matching results are unchanged, and upgrading needs no migration window because every reader falls back to the inline blobs. Completing the split is a separate, manual step: run./migrate.sh, see Maintenance. NOTE: the in-place migration temporarily needs about as much free disk as your stored disassembly - +15 GB on an 11.6M function corpus - and does not hand it back until you compact the collection. NOTE: once migrated, older MCRIT versions can no longer read the database and fail quietly - they start up healthy and then behave as if no function had disassembly - so run the migration's--mode unsplitbefore ever downgrading the images. Also NOTE: theconfig/copies shipped here had drifted behind MCRIT and were missing the settings introduced in 1.6.0/1.6.1, which silently left vectorised matching, numpy candidate accumulation, the persistent MatchingCache, concurrent signature fetch and the candidate-pair budget all disabled; they are regenerated from 1.7.0 stock in this bump, so matching should get noticeably faster with unchanged results - if you have editedconfig/locally, merge rather than overwrite.), MCRITweb 1.4.8 - 2026-08-20: MCRIT 1.6.2 (reliability release: the worker survives a mongod restart instead of exiting, a dead child process fails its job instead of reporting it finished, queue polling and the sha256/family/function-name lookups are served by indexes, and matching memory is bounded by a candidate-pair budget - see docs/TUNING.md; matching results are unchanged), MCRITweb 1.4.8 (NOTE: the first start after this bump builds two new indexes on the
functionscollection, which blocks for minutes on a multi-million function corpus - plan the restart accordingly) - 2026-08-11: MCRIT 1.6.1 (the 1.6.0 optimisations are now on by default — matching runs single-process and the MatchingCache persists under a 512 MiB budget; results are unchanged, see docs/TUNING.md), MCRITweb 1.4.8 (the session cookie is now
Secure— an instance served over plain HTTP needsSESSION_COOKIE_SECURE = Falseininstance/config.pyor logins fail; the NGINX-terminated deployment here is unaffected) - 2026-08-11: MCRIT 1.6.0 (major matching-performance release, up to 4.4x with the new opt-in optimisations — all default off, see docs/TUNING.md), MCRITweb 1.4.7
- 2026-08-07: MCRIT 1.5.3, MCRITweb 1.4.7 (Flask 3 / Werkzeug 3 — rebuild the mcritweb image, a code-only pull keeps the old Flask)
- 2026-08-06: MCRIT 1.5.3, MCRITweb 1.4.6 (overall code quality and security improvements)
- 2026-08-04: MCRIT 1.5.3 (~7x faster matching report loading), MCRITweb 1.4.2
- 2026-08-04: MCRIT 1.5.2 (Dalvik capability, shingler packaging fix), MCRITweb 1.4.1
- 2026-07-16: MCRIT 1.5.0 (worker+server moved to ubuntu24.04 / python3.12), MCRITweb 1.4.1
- 2025-12-10: MCRIT 1.4.3, MCRITweb 1.4.1
- 2025-12-08: MCRIT 1.4.3, MCRITweb 1.4.0
- 2025-08-22: MCRIT 1.4.1, MCRITweb 1.3.6