Schema generator
ClinVar SQLite Schema Generator.
This module provides schema generation for ClinVar databases: - RCV (Reference ClinVar Assertion) - condition-centric format - VCV (Variant Call Variation) - variant-centric format
The SQL schemas are based on the following XSD: - RCV: ClinVar_RCV_weekly.xsd v2.2 (August 6, 2025) - VCV: ClinVar_VCV.xsd v2.5 (August 6, 2025)
Notes
This module creates only the database schema structure (empty tables). To populate databases with actual ClinVar data, use clinvar_parser.py.
- class clinvar_build.schema_generator.ClinVarSchemaGenerator[source]
Builder for ClinVar SQLite database schemas.
This class provides methods to create database schemas for ClinVar RCV (condition-centric), and VCV (variant-centric), It handles connecting to the SQLite database, executing table and index creation, and safely closing the connection.
- conn
Active SQLite connection, or None if not connected.
- Type:
sqlite3.Connection or None
- cursor
Cursor for executing SQL statements, or None if not connected.
- Type:
sqlite3.Cursor or None
- create_schema(db_path, db_config, db_indices=None, name=None)[source]
Create a SQLite schema for a given ClinVar database.
- __str__() str[source]
Return human-readable string representation.
- Returns:
Human-readable description of the builder
- Return type:
- create_schema(db_path: str | Path, db_config: list[str], db_indices: list[str] | None, name: str | None = None) None[source]
Create a SQLite schema for ClinVar (RCV, VCV).
- Parameters:
db_path (str or Path) – Path to SQLite database file.
db_config (list [str]) – List of SQL CREATE TABLE statements.
db_indices (list [str] or None, default None) – List of SQL CREATE INDEX statements.
name (str or None, default None) – Optional name for logging. Defaults to the filename stem of db_path.
- Return type:
None
- class clinvar_build.schema_generator.JSONSchemaLoader[source]
Loader that derives SQLite Data Definition Language (DDL) from a parsed JSON config file.
Reads a parse JSON config file and converts entity definitions, constraint annotations, and index specifications into
CREATE TABLEandCREATE INDEXSQL strings suitable forClinVarSchemaGenerator.create_schema().- tables
Complete
CREATE TABLESQL statements. Populated afterload()is called.- Type:
list [str] or None
- indexes
Complete
CREATE INDEXSQL statements. Populated afterload()is called.- Type:
list [str] or None
- Class Attributes
- ----------------
- CAST_TYPE_MAP
Mapping from parse JSON
castvalues to SQLite column types.- Type:
dict [str, str]
Examples
>>> from pathlib import Path >>> loader = JSONSchemaLoader() >>> tables, indexes = loader.load(Path("resources/.../rcv_parse.json")) >>> tables # list of CREATE TABLE statements >>> indexes # list of CREATE INDEX statements
- __weakref__
list of weak references to the object
- load(path: Path) tuple[list[str], list[str]][source]
Load schema Data Definition Language (DDL) from a parse JSON config file.
Derives table and index DDL from the entity definitions and constraint annotations in the JSON, returning lists in the same format expected by
ClinVarSchemaGenerator.create_schema().- Parameters:
path (Path) – Path to a parse JSON config file (
*_parse.json).- Returns:
tables (list [str]) – Complete
CREATE TABLESQL statements.indexes (list [str]) – Complete
CREATE INDEXSQL statements.
Examples
>>> from pathlib import Path >>> loader = JSONSchemaLoader() >>> tables, indexes = loader.load(Path("rcv_parse.json"))
Parser
ClinVar XML to SQLite database parser
This module provides a comprehensive parser for converting ClinVar XML files into SQLite databases. It implements a configuration-driven architecture that supports both RCV (Reference ClinVar) and VCV (Variation ClinVar) XML formats through JSON-based table and column specifications.
The parser is implemented using iterative parsing and batch commits, minimising memory usage. Progress tracking is enabled through a SQL table, allowing for restarts after the last committed record.
- class clinvar_build.parser.ClinVarParser(config: dict[str, Any] | None = None)[source]
Parser for ClinVar XML files.
- Parameters:
config (dict[str, Any]) – A dictionary with instructions to parser an XML file to a SQLite database.
- conn
Active SQLite connection, or None if not connected.
- Type:
sqlite3.Connection or None
- cursor
Cursor for executing SQL statements, or None if not connected.
- Type:
sqlite3.Cursor or None
- stats
Statistics on parsed records
- Type:
dict [str, int]
- Class Attributes
- ----------------
- _SQL_COLS_PATTERN
Compiled regex to extract column names from SQL INSERT statements
- Type:
- parse_file(xml_path: str | Path, db_path: str | Path, batch_size: int = 10, validate_foreign_keys: bool = True, enforce_foreign_keys: bool = False, xsd_path: str | Path | None = None, xsd_strict: bool = False, resume: bool = True, count_duplicates: bool = False) None[source]
Parse ClinVar XML file into SQLite database.
- Parameters:
xml_path (str or Path) – Path to ClinVar XML file
db_path (str or Path) – Path to SQLite database file
batch_size (int, default 10) – Number of top-level records (including children) to commit at once.
validate_foreign_keys (bool, default True) – Whether to validate foreign key integrity after parsing
enforce_foreign_keys (bool, default False) – Whether to enforce foreign keys during parsing (slower but catches errors immediately). If False, foreign keys are only validated after parsing completes.
xsd_path (str, Path, or None, default None) – Optional path to XSD schema for XML validation
xsd_strict (bool, default False) – If False, allows XML elements not defined in XSD
resume (bool, default True) – If True, attempt to resume from last checkpoint. If no checkpoint exists, starts fresh. If False, always starts fresh (existing data may cause constraint violations).
count_duplicates (bool, default False) – After building the database print potential duplicated rows per table
Queries
ClinVar VCV/RCV database queries.
This module encodes some limited, but frequently used, queries to extract clinvar data.
- class clinvar_build.queries.BaseQuery(db_path: str | Path, secondary_db_path: str | Path | None = None)[source]
Abstract base class for ClinVar database queries.
Handles database connection, parameter building, and query execution. Subclasses implement specific query logic.
- Parameters:
db_path (str or Path) – Path to the primary SQLite database file.
secondary_db_path (str or Path, optional) – Path to secondary database for enrichment.
- __exit__(*_: object) None[source]
Exit context manager and close connections.
Notes
Exception parameters are part of the context manager protocol but are not used in this base implementation. Connections are closed regardless of whether an exception occurred, and any exception propagates normally. Subclasses may override to add custom exception handling such as transaction rollback.
- __init__(db_path: str | Path, secondary_db_path: str | Path | None = None)[source]
Initialize querier.
- __weakref__
list of weak references to the object
- class clinvar_build.queries.FetchRangeQuery(db_path: str | Path, secondary_db_path: str | Path | None = None)[source]
Fetch variants within genomic ranges.
This query class operates on the VCV (Variation Archive) database as its primary source. When enrichment is enabled, condition data is fetched from the RCV database and merged via canonical_spdi.
Assembly behaviour
Returns one row per (variant × assembly). Each
GenomicRangecarries an assembly (default'GRCh38'); only variants whose coordinates in that assembly fall within the range are returned. Usegrch38andgrch37boolean params to restrict which assemblies are joined.- query(ranges, genes, classifications, ...)[source]
Fetch variants within specified genomic regions.
Examples
Single range (GRCh38, the default): >>> q = FetchRangeQuery(“clinvar_vcv.db”, “clinvar_rcv.db”) >>> df = q.query(ranges=[GenomicRange(“19”, 11000000, 12000000)])
GRCh37 range: >>> df = q.query(ranges=[GenomicRange(“19”, 10850000, 11850000, “GRCh37”)])
Both assemblies for the same locus: >>> df = q.query(ranges=[ … GenomicRange(“19”, 11000000, 12000000, “GRCh38”), … GenomicRange(“19”, 10850000, 11850000, “GRCh37”), … ])
With filters and enrichment: … ranges=[GenomicRange(“19”, 11000000, 12000000)], … genes=[“LDLR”], … classifications=[“Pathogenic”], … enrich=True, … )
- query(ranges: list[GenomicRange], genes: list[str] | None = None, classifications: list[str] | None = None, grch38: bool = True, grch37: bool = True, enrich: bool = False, output_path: str | Path | None = None) DataFrame | int[source]
Fetch variants within genomic ranges.
- Parameters:
ranges (list [GenomicRange]) – Genomic ranges to query. Each
GenomicRangecarries an assembly (default'GRCh38'); only coordinates for that assembly are matched. Pass ranges for multiple assemblies to retrieve results across both GRCh37 and GRCh38.genes (list [str], default None) – Filter by gene symbols.
classifications (list [str], default None) – Filter by classification.
grch38 (bool, default True) – Include GRCh38 assembly rows in the JOIN.
grch37 (bool, default True) – Include GRCh37 assembly rows in the JOIN.
enrich (bool, default False) – Enrich with RCV condition data.
output_path (str or Path, default None) – If provided, stream results to file and return row count.
- Returns:
Variants within specified ranges. One row per variant-assembly combination, or row count when streaming to disk.
- Return type:
pd.DataFrame or int
- Raises:
ValueError – If both
grch38andgrch37are False.
- class clinvar_build.queries.FetchVariantQuery(db_path: str | Path, secondary_db_path: str | Path | None = None)[source]
Fetch specific variants by identifier, name, or position.
This query class operates on the VCV (Variation Archive) database as its primary source. When enrichment is enabled, condition data is fetched from the RCV database and merged via canonical_spdi.
Assembly behaviour
Returns one row per (variant × assembly). When
locationis given, only the assembly specified inlocation.assembly(default GRCh38) is returned.- query(vcv_accessions, variation_ids, variant_names, ...)[source]
Fetch variants matching specified identifiers or positions.
Examples
>>> q = FetchVariantQuery("clinvar_vcv.db", "clinvar_rcv.db") >>> # two rows per variant (GRCh38 + GRCh37) >>> df = q.query(vcv_accessions=["VCV000017639"]) >>> # GRCh37 coordinates only >>> df = q.query(location=GenomicLocation("19", 11089362), grch38=False)
- query(vcv_accessions: list[str] | None = None, variation_identifiers: list[int] | None = None, variant_names: list[str] | None = None, variant_name_pattern: str | None = None, location: GenomicLocation | None = None, spdi: str | None = None, grch38: bool = True, grch37: bool = True, enrich: bool = False, output_path: str | Path | None = None) DataFrame | int[source]
Fetch variants by identifier or position.
- Parameters:
vcv_accessions (list [str], default None) – VCV accession numbers (e.g., [‘VCV000017639’]).
variation_identifiers (list [int], default None) – ClinVar variation IDs (e.g., [17639]).
variant_names (list [str], default None) – Exact variant names (e.g., [‘NM_000527.5:c.1A>G’]).
variant_name_pattern (str, default None) – Variant name pattern using SQL LIKE patterns.
location (GenomicLocation, default None) – Genomic location as (chr, position, ref, alt). Only chr and position are required.
location.assembly(default ‘GRCh38’) restricts results to that assembly.spdi (str, default None) – Canonical SPDI notation.
grch38 (bool, default True) – Include GRCh38 assembly rows.
grch37 (bool, default True) – Include GRCh37 assembly rows.
enrich (bool, default False) – Enrich with RCV condition data.
output_path (str or Path, default None) – If provided, stream results to file and return row count.
- Returns:
Matching variants. One row per variant-assembly combination, or row count when streaming to disk.
- Return type:
pd.DataFrame or int
- Raises:
ValueError – If both
grch38andgrch37are False, or if thelocationassembly conflicts with the assembly flags.
- class clinvar_build.queries.FilterQuery(db_path: str | Path, secondary_db_path: str | Path | None = None)[source]
Filter variants by gene, condition, and classification.
This query class operates on the VCV (Variation Archive) database as its primary source. When enrichment is enabled, condition data is fetched from the RCV database and merged via canonical_spdi.
Returns one row per variant-assembly-condition combination. Each row contains QONames.assembly, QONames.chrom, QONames.start, QONames.stop, QONames.position_vcf, QONames.reference_allele_vcf, and QONames.alternate_allele_vcf columns for that assembly. Because QONames.cc_name, QONames.cc_db, QONames.cc_condition_id, and QONames.rc_classification are one-to-many with respect to a variant, a single variant can produce multiple rows.
- query(genes, condition_patterns, condition_ids, classifications, ...)[source]
Filter variants matching specified criteria.
Examples
>>> q = FilterQuery("clinvar_vcv.db", secondary_db_path="clinvar_rcv.db") >>> df = q.query( ... genes=["LDLR"], ... classifications=["Pathogenic"], ... enrich=True, ... )
- query(genes: list[str] | None = None, condition_names: list[str] | None = None, condition_identifiers: list[tuple[str, str]] | None = None, classifications: list[str] | None = None, grch38: bool = True, grch37: bool = True, enrich: bool = False, output_path: str | Path | None = None) DataFrame | int[source]
Filter variants with optional enrichment.
- Parameters:
genes (list [str], default None) – Filter by gene symbols (e.g., [‘LDLR’, ‘APOB’]).
condition_names (list [str], default None) – Filter by condition name using SQL LIKE patterns (e.g., [‘%cardiomyopathy%’]) or exact matches.
condition_identifiers (list [tuple], default None) – Filter by condition identifier as (db, id) tuples (e.g., [(‘MedGen’, ‘C0007194’), (‘OMIM’, ‘115200’)]).
classifications (list [str], default None) – Filter by classification (e.g., [‘Pathogenic’]).
grch38 (bool, default True) – Include GRCh38 assembly rows.
grch37 (bool, default True) – Include GRCh37 assembly rows.
enrich (bool, default False) – Enrich with data from secondary database.
output_path (str or Path, default None) – If provided, stream results to file and return row count.
- Returns:
Filtered variants. One row per variant-assembly-condition combination, or row count when streaming to disk.
- Return type:
pd.DataFrame or int
- Raises:
ValueError – If both
grch38andgrch37are False.
- class clinvar_build.queries.GenomicLocation(chr: str, position: int, ref: str | None = None, alt: str | None = None, assembly: Assembly = 'GRCh38')[source]
Genomic location for variant lookup.
- Parameters:
chr (str) – Chromosome (e.g., ‘19’, ‘X’).
position (int) – VCF position (1-based).
ref (str, default None) – Reference allele.
alt (str, default None) – Alternate allele.
assembly ({‘GRCh37’, ‘GRCh38’}, default GRCh38) – Genome assembly.
Examples
>>> loc = GenomicLocation("19", 11089362) >>> loc = GenomicLocation("19", 11089362, "A", "G") >>> loc = GenomicLocation("19", 11089362, assembly="GRCh37")
- __getnewargs__()
Return self as a plain tuple. Used by copy and pickle.
- static __new__(_cls, chr: str, position: int, ref: str | None = None, alt: str | None = None, assembly: Assembly = 'GRCh38')
Create new instance of GenomicLocation(chr, position, ref, alt, assembly)
- __replace__(**kwds)
Return a new GenomicLocation object replacing specified fields with new values
- __repr__()
Return a nicely formatted representation string
- class clinvar_build.queries.GenomicRange(chr: str, start: int, end: int, assembly: Assembly = 'GRCh38')[source]
Genomic range for variant queries.
- Parameters:
chr ('str') – Chromosome (e.g., ‘19’, ‘X’).
start ('int') – Start position as basepair position.
end ('int') – End position as basepair position.
assembly ({'GRCh37', 'GRCh38'}, default 'GRCh38') – Genome assembly.
Examples
>>> r = GenomicRange("19", 11000000, 12000000) >>> r = GenomicRange("19", 11000000, 12000000, assembly="GRCh37")
- __getnewargs__()
Return self as a plain tuple. Used by copy and pickle.
- static __new__(_cls, chr: str, start: int, end: int, assembly: Assembly = 'GRCh38')
Create new instance of GenomicRange(chr, start, end, assembly)
- __replace__(**kwds)
Return a new GenomicRange object replacing specified fields with new values
- __repr__()
Return a nicely formatted representation string
- class clinvar_build.queries.ListConditionsQuery(db_path: str | Path, secondary_db_path: str | Path | None = None)[source]
List condition/disease identifiers and metadata.
This query class operates on the RCV database as its primary source.
Examples
List available sources: >>> q = ListConditionsQuery(“clinvar_rcv.db”) >>> df = q.available_sources()
Query conditions from a specific source: >>> df = q.query(source=”MedGen”)
Include variant frequency: >>> df = q.query( … source=”MedGen”, … include_frequency=True, … min_frequency=10, … )
- available_sources() DataFrame[source]
List available condition identifier sources from RCV database.
- Returns:
Sources with counts.
- Return type:
pd.DataFrame or int
- query(source: str | None = None, include_frequency: bool = False, min_frequency: int = 1, output_path: str | Path | None = None) DataFrame | int[source]
Get unique condition identifiers from RCV database.
- Parameters:
source ('str', default None) – Filter by source database (e.g., ‘MedGen’, ‘OMIM’). If None, returns all sources.
include_frequency ('bool', default False) – Include variant count per condition.
min_frequency ('int') – Minimum variant count to include (default 1). Only applies when include_frequency is True.
output_path ('str' or 'Path', default None) – If provided, stream results to file and return row count.
- Returns:
Unique conditions with identifiers, or row count if streaming to disk.
- Return type:
pd.DataFrame or int
General utils
General utility functions for the clinvar-build module
- clinvar_build.utils.general.assign_empty_default(arguments: list[Any], empty_object: Callable[[], Any]) list[Any][source]
Takes a list of arguments, checks if these are NoneType and if so assigns them ‘empty_object’.
- Parameters:
arguments (list [any]) – A list of arguments which may be set to NoneType.
empty_object (Callable)
object (A function that returns a mutable) – Examples include a list or a dict.
- Returns:
new_arguments – List with NoneType replaced by empty mutable object.
- Return type:
Examples
>>> assign_empty_default(['hi', None, 'hello'], empty_object=list) ['hi', [], 'hello']
Notes
This function helps deal with the pitfall of assigning an empty mutable object as a default function argument, which would persist through multiple function calls, leading to unexpected/undesired behaviours.
Configuration tools
Configuration parsing and XML validation utilities for ClinVar Build.
This module provides utilities for parsing configuration files, and managing logging output for long-running operations. It includes classes for handling block-based configuration files and property management with controlled access.
- class clinvar_build.utils.config_tools.ManagedProperty(name: str, types: tuple[type] | type | None = None)[source]
A generic property factory defining setters and getters, with optional type validation.
- Parameters:
name (str) – The name of the setters and getters
types (Type, default NoneType) – Either a single type, or a tuple of types to test against.
- set_with_setter(instance, value)[source]
Enables the setter, sets the property value, and then disables the setter, ensuring controlled updates.
- Returns:
A property object with getter and setter.
- Return type:
- class clinvar_build.utils.config_tools.ProgressHandler(stream=None)[source]
Custom handler that updates progress in place.
Uses ANSI escape codes to overwrite previous output instead of printing new lines. Useful for progress updates during long-running operations.
- Parameters:
stream (file-like object, optional) – Output stream. Defaults to sys.stdout.
- _last_line_count
Number of lines in the previous message, used to calculate how far to move the cursor up.
- Type:
int
Examples
>>> progress_logger = logging.getLogger('progress') >>> handler = ProgressHandler() >>> handler.setFormatter(logging.Formatter('%(asctime)s - %(message)s')) >>> progress_logger.addHandler(handler) >>> progress_logger.info("Processing: 100 records") >>> progress_logger.info("Processing: 200 records") # Overwrites previous
- clinvar_build.utils.config_tools.check_environ(environ_variable: str = 'CLINVAR_BUILD_CONFIG', fall_back: str | Path | None = PosixPath('/usr/local/share/clinvar_build/clinvar_config')) str[source]
Retrieve an environment variable pointing to a directory path, with optional fallback path.
Attempts to retrieve the specified environment variable. If the variable is not set, the function will attempt to use the fallback path if provided. This is useful for configuration management where environment variables may not always be explicitly set.
- Parameters:
environ_variable (str) – The name of the environment variable to retrieve.
fall_back (str, Path or None) – A fallback path to return if the environment variable is not set. If None, an error will be raised when the environment variable is missing.
- Returns:
A directory path.
- Return type:
- Raises:
Notes
The function will not check whether the path is available or whether permissions allow for read or write access
Warning
- UserWarning
Issued when the environment variable is not set and the fallback path is used instead.
Examples
>>> check_environ("MY_VAR") '/path/to/default/config'
Parser tools
XML parsing and SQLite database utilities for ClinVar Build.
This module provides classes and functions for parsing large XML clinvar files and loading them into SQLite databases. It includes utilities for database validation, progress tracking, error formatting, and XML inspection.
- class clinvar_build.utils.parser_tools.SQLiteParser(config: dict[str, Any] | None = None)[source]
A general SQLite parser.
- Parameters:
config (dict[str, Any]) – A dictionary with instructions to parser an XML file to a SQLite database.
- conn
Active SQLite connection, or None if not connected.
- Type:
sqlite3.Connection or None
- cursor
Cursor for executing SQL statements, or None if not connected.
- Type:
sqlite3.Cursor or None
- stats
Statistics on parsed records.
- Type:
dict [str, int]
- config
A configuration dictionary with parsing instructions.
- Type:
dict [str, any]
- __repr__() str[source]
Return unambiguous string representation.
- Returns:
String representation suitable for debugging
- Return type:
- __str__() str[source]
Return human-readable string representation.
- Returns:
Human-readable description of the parser
- Return type:
- count_duplicates()[source]
Count duplicate rows in all tables based on non-id columns.
For each table in the configuration, identifies duplicate rows where all columns (except the primary key ‘id’) are identical, treating NULL values as equal.
Notes
Uses SQLite’s rowid to identify duplicates, counting rows that would be removed (keeping the row with lowest rowid). NULL values are treated as equal via GROUP BY.
- validate_database() dict[str, Any][source]
Run comprehensive database validation checks.
Performs multiple validation checks including foreign key integrity, table row counts, and basic statistics.
- Returns:
Dictionary containing validation results and statistics
- Return type:
- Raises:
ValueError – If foreign key violations are found
Examples
>>> with parser._connection(db_path): ... results = parser.validate_database() >>> print(results) {'foreign_keys': 'valid', 'total_records': 12345, ...}
- class clinvar_build.utils.parser_tools.ViewXML(xml_path: str | Path, tag_name: str, index: int = 0)[source]
Class to view and inspect XML records.
This class loads a single XML element and provides utilities for inspecting and navigating its structure, particularly useful for large ClinVar XML files.
- Parameters:
xml_path (str or Path) – Path to ClinVar XML file (supports .gz compression)
tag_name (str) – XML tag name to search for (e.g., ‘VariationArchive’)
index (int, default 0) – Which occurrence of the tag to load (0-indexed)
- xml_path
Path to the source XML file
- Type:
Path
- tag_name
The tag name that was searched for
- Type:
str
- index
The index of the loaded element
- Type:
int
- element
The loaded XML element, or None if not found
- Type:
ET.Element or None
Examples
>>> viewer = ViewXML('clinvar.xml.gz', 'VariationArchive', index=5) >>> viewer.show_tree() >>> paths = viewer.find_all_paths()
- __init__(xml_path: str | Path, tag_name: str, index: int = 0)[source]
Initialise ViewXML with a single XML element.
- __repr__() str[source]
Return unambiguous string representation.
- Returns:
String that could recreate the object
- Return type:
- __str__() str[source]
Return human-readable string representation.
- Returns:
Summary of the loaded element
- Return type:
- count_all_tags() dict[str, int][source]
Count all descendant tags in loaded element tree.
- Returns:
dict of {str – Dictionary mapping tag names to their occurrence counts, ordered by frequency (most common first)
- Return type:
int}
- Raises:
ValueError – If no element is loaded
- find_all_paths() list[str][source]
Find all unique XPath-like paths in the loaded element tree.
- Returns:
Sorted list of all unique paths found in the tree
- Return type:
- Raises:
ValueError – If no element is loaded
Examples
>>> viewer = ViewXML('clinvar.xml.gz', 'VariationArchive') >>> paths = viewer.find_all_paths() >>> print(paths[:5])
- show_children(indent: int = 0) None[source]
Display immediate children of loaded element with summary info.
- Parameters:
indent (int, default 0) – Indentation level for output formatting
- Raises:
ValueError – If no element is loaded
- show_tree(max_depth=4, indent=0)[source]
” Display element tree structure up to specified depth.
- Parameters:
max_depth (int, default 4) – Maximum depth to display
indent (int, default 0) – Starting indentation level
- Raises:
ValueError – If no element is loaded
- clinvar_build.utils.parser_tools.configure_logging(verbosity: int) None[source]
Configure logging based on verbosity level.
This function configures the root logger, which affects all loggers in the application through inheritance.
- Parameters:
verbosity (int) – Number of -v flags. 0 = WARNING, 1 = INFO, 2 = DEBUG, 3 = TRACE
- clinvar_build.utils.parser_tools.open_xml_file(xml_path: str | Path, xsd_path: str | Path | None = None, strict: bool = True, verbose: bool = True)[source]
Helper function to open an XML file, handling compression based on extension.
Automatically closes the file handle when exiting the context, even if an exception occurs.
- Parameters:
xml_path (str) – The path to the file.
xsd_path (str, Path, or None, default None) – Optional path to XSD schema file for validation. If provided, validates XML before yielding the file handle.
strict (bool, default True) – If False, ignores elements in the XML that are not in the XSD. Only used when xsd_path is provided.
verbose (bool, default True) – If True, warns about validation issues when strict=False. Only used when xsd_path is provided.
- Yields:
file-like object – An open file handle for the compressed or uncompressed XML.
- Raises:
FileNotFoundError – If xml_path or xsd_path does not exist
IOError – If file cannot be opened
XMLValidationError – If XSD validation fails (when xsd_path is provided)
- clinvar_build.utils.parser_tools.trace(self, message, *args, **kwargs)[source]
custom trace method for detailed logging.
- clinvar_build.utils.parser_tools.validate_xml(xml_path: str | Path, xsd_path: str | Path, strict: bool = True, verbose: bool = True) _ElementTree[source]
Validates an XML file against an XSD schema.
- Parameters:
xml_path (str or Path) – Path to the XML file.
xsd_path (str or Path) – Path to the XSD file.
strict (bool, default True) – If False, ignores elements in the XML that are not in the XSD.
- Returns:
The parsed XML document.
- Return type:
etree._ElementTree
- Raises:
XMLValidationError – Raised if the XSD and XML are incompatible.