logoalt Hacker News

ltbarcly3today at 4:40 PM4 repliesview on HN

Our db clearing takes like 5ms (large schema from mature company, not a toy). We start by restoring a production schema dump which ensures our test db / devdb schema is basically identical to what we run in production. Any migrations you are working on in your branch get added to the restore of the prod schema after it runs. Building a schema from ORM definitions is what you do if you don't care about your life or time. Clearing the test db takes something similar. This is fine, using template db's is a good way to make a scratch copy of a local database to test migrations (rsync is better if you are technical enough to use it to restore after destructive changes).

Timings:

  - create testdb: 9ms
  - restore prod schema: 500ms (done once per test process)
  - clear test data in 96 tables between tests that write to db (5ms)
The fastest way to clear a test db is to run a query to get every schema/table name, then run ";".join("`DELETE FROM {schema}.{tablename};" for schema, table in my_tables) after putting the db in replica mode. This takes single digit ms a lot of the time even with a decent amount of test data. I've done it every which way and this is by far the fastest way to clear data between tests.

    -- This will work on basically any postgresql database with basically any schema so just use it.
    test_db_2235191=# CREATE OR REPLACE PROCEDURE public.delete_all_table_data()
        LANGUAGE plpgsql
        AS $procedure$
        DECLARE
            target record;
            previous_replication_role text;
        BEGIN
            previous_replication_role :=
                current_setting('session_replication_role');
        PERFORM set_config('session_replication_role', 'replica', true);

        BEGIN
            FOR target IN
                SELECT namespace.nspname AS schema_name,
                       relation.relname AS table_name
                FROM pg_catalog.pg_class AS relation
                JOIN pg_catalog.pg_namespace AS namespace
                  ON namespace.oid = relation.relnamespace
                WHERE relation.relkind = 'r'
                  AND namespace.nspname NOT LIKE 'pg\_%' ESCAPE '\'
                  AND namespace.nspname <> 'information_schema'
                ORDER BY namespace.nspname, relation.relname
            LOOP
                RAISE NOTICE 'Deleting %.%',
                    target.schema_name,
                    target.table_name;

                EXECUTE format(
                    'DELETE FROM %I.%I',
                    target.schema_name,
                    target.table_name
                );
            END LOOP;
        EXCEPTION
            WHEN OTHERS THEN
                PERFORM set_config(
                    'session_replication_role',
                    previous_replication_role,
                    true
                );
                RAISE;
        END;

        PERFORM set_config(
            'session_replication_role',
            previous_replication_role,
            true
        );
    END;
    $procedure$;
    CREATE PROCEDURE
    Time: 0.840 ms


    test_db_2235191=# CALL public.delete_all_table_data();
    NOTICE:  ... (notices removed for 96 tables)
    CALL
    Time: 5.855 ms

Replies

stephentoday at 7:07 PM

We do similar, although lean into our strict "every table as a sequence" and "all FKs are deferred" conventions and only issue DELETEs for tables that actually were inserted by the test

https://github.com/joist-orm/joist-orm/blob/16cc73f148b6f962...

I forgot the speedup this got us on a 400-500 table schema, but it was noticeable -- curious if you could do the same / what the perf impact would be.

leontrolskitoday at 6:20 PM

I concur with this approach. TRANSACTION-y tests (the default in Django) often don't quite line up with reality and make it hard to eg. drop in a breakpoint and run a server against the test's db state.

I've experimented (see below) with TEMPLATE dbs and such in Python (with inspiration from this library). IMHO the "around 100ms" mark is pretty slow for a big test suite. Interestingly, pg_restore is only twice as slow as TEMPLATEs.

https://github.com/leontrolski/postgresql-testing

I'd be interested about how all this compares to snapshotting the postrgres dir with ZFS and restoring to that, but don't have a Linux box to hand.

brandurtoday at 5:53 PM

Interesting. How many database copies do you bring up when the test suite starts running, and how is parallelism handled?

show 1 reply
leontrolskitoday at 6:21 PM

Being to lazy to think or test it - does the above reset SEQUENCEs?

show 1 reply