September 28, 2026 · Matt Irvin
Apache Hop is the open-source destination TOS shops keep asking about after Open Studio's end of life. Here is what actually ports, what has to be rewritten by hand, how tMap and Java routines translate, and when Hop is the wrong answer.
Read moreSeptember 21, 2026 · Matt Irvin
A practical pattern for moving clinical data with Talend: parsing HL7 v2 segments into a normalized staging model, calling FHIR R4 REST APIs with pagination and OAuth, pulling NDJSON from $export, and keeping PHI handling defensible.
Read moreSeptember 21, 2026 · Matt Irvin
TAC and JobServers are the last on-prem piece most Talend shops are still running. Here is the inventory-first runbook we use to move projects, schedules, execution plans, contexts, and users onto Talend Management Console and Remote Engines without losing a night of loads.
Read moreSeptember 18, 2026 · Matt Irvin
Green jobs do not mean good data. How to profile a source, store quality rules as data, evaluate them with one generic Talend job, keep freshness and volume checks, and gate the publish step so bad batches never reach your users.
Read moreSeptember 17, 2026 · Matt Irvin
BigQuery offers four write paths and Talend can use all of them — only one belongs in a nightly batch. Choosing load jobs over GCS, authenticating without a stray JSON key, pinning NUMERIC and timestamp types, and writing a MERGE that prunes partitions instead of scanning the whole table.
Read moreSeptember 15, 2026 · Matt Irvin
Most warehouse overspend is created upstream, in the integration layer. How to attribute Snowflake, BigQuery, and Databricks cost back to individual Talend jobs with query tags, find the handful of jobs burning most of the bill, and fix them with batching, incremental loads, and warehouse discipline.
Read moreSeptember 14, 2026 · Matt Irvin
Auditors and platform teams now expect lineage for every pipeline, and Talend jobs are usually the black box in the middle. Here is how we emit OpenLineage run events from Talend, model job-to-job dependencies, add column-level mappings from tMap, and land it all in Marquez or a catalog like DataHub.
Read moreSeptember 11, 2026 · Matt Irvin
Retrieval-augmented generation is only as good as the pipeline that feeds it. A practical Talend design for crawling source documents, chunking text, calling an embeddings API in batches, and upserting vectors into Postgres pgvector idempotently — with re-embedding, cost control, and deletion handling.
Read moreSeptember 10, 2026 · Matt Irvin
The pattern we use to land Talend output in Databricks: Parquet staging in cloud storage, external locations and volumes governed by Unity Catalog, COPY INTO for the append case, and an idempotent MERGE for upserts. Plus the JDBC, auth, and cost mistakes that show up on every first cutover.
Read moreSeptember 9, 2026 · Matt Irvin
Talend has no tIcebergOutput component, and it does not need one. The three patterns we deploy for landing data in Iceberg tables, replay-safe MERGE logic, schema evolution rules, and the compaction and snapshot-expiry jobs teams forget until queries slow down.
Read moreSeptember 9, 2026 · Matt Irvin
Most Talend shops deploy on hope and a smoke test. Here is a practical testing stack: Studio test cases for component logic, SQL-based data assertions for outputs, and a golden-dataset regression suite you can run in CI before every release.
Read moreSeptember 9, 2026 · Matt Irvin
A practical tuning walkthrough for slow Talend jobs: how to find where the time actually goes, fix tMap lookup memory, use parallelism properly, replace row-by-row writes with bulk loaders, and push work down to the warehouse.
Read moreSeptember 9, 2026 · Matt Irvin
A practical pattern for catching every Talend failure once: tLogCatcher/tStatCatcher wired into a reusable joblet, a job-run table you can query, JSON logs your platform team can ingest, and alerts that fire on the runs that never started.
Read moreSeptember 9, 2026 · Matt Irvin
Talend lands the data and dbt models it — but the seam between them is usually a cron gap and a prayer. A practical guide to load-audit tables, dbt freshness and completeness tests, three ways to trigger dbt from Talend, and which transformations belong in which tool.
Read moreSeptember 9, 2026 · Matt Irvin
How to build Kafka consumers and producers in Talend: tKafkaInput/tKafkaOutput configuration, consumer groups and offset commits, why 'exactly-once' is really idempotent-once-in-the-warehouse, schema handling, and when a long-running streaming job is the wrong answer.
Read moreSeptember 9, 2026 · Matt Irvin
A working recipe for Type 2 dimension loads in Talend: tMap-based hash comparison versus tDBSCD, choosing a change-detection hash, handling late-arriving and out-of-order rows, making reruns idempotent, and the warehouse MERGE that closes and opens versions in one pass.
Read moreSeptember 9, 2026 · Matt Irvin
Moving from on-prem JobServers to Qlik Talend Cloud means someone still has to run the engine. A practical guide to Remote Engine sizing, outbound-only networking, secrets, engine pools, and the runtime patterns that keep hybrid pipelines predictable.
Read moreSeptember 9, 2026 · Matt Irvin
How to build a repeatable Talend masking pipeline: choosing between substitution, format-preserving encryption and shuffling, keeping joins intact across tables, handling free-text and dates, and proving the output is actually de-identified.
Read moreSeptember 9, 2026 · Matt Irvin
A practical pattern for landing Talend output in Microsoft Fabric: writing Parquet to OneLake over the ADLS Gen2 endpoint with a service principal, registering Lakehouse tables, and using COPY INTO for Warehouse loads with idempotent reruns.
Read moreSeptember 9, 2026 · Matt Irvin
Source systems add, rename, and retype columns without warning. This tutorial shows how to survive it in Talend: when to use a Dynamic schema with tSetDynamicSchema, how to detect drift before it corrupts a load, how to auto-evolve a Snowflake or Postgres target safely, and how to turn all of it into a data contract your upstream team actually has to honour.
Read moreSeptember 9, 2026 · Matt Irvin
A working pattern for deduplicating customer and supplier data in Talend: standardize first, block to keep the comparison count sane, tune tMatchGroup thresholds against a labelled sample, then apply explicit survivorship rules and keep a crosswalk of every merge you make.
Read moreSeptember 9, 2026 · Matt Irvin
How to package a Talend standalone job as a Docker image and run it on Kubernetes: a boring Dockerfile, context values as arguments, secrets from the cluster store, CronJob settings that prevent overlapping loads, JVM memory limits that avoid silent OOM kills, and exit codes that actually fail.
Read moreAugust 25, 2026 · Matt Irvin
Qlik retired Talend Open Studio on January 31, 2024. Two years on, here is what still works, what is quietly rotting, your four realistic paths off it with honest effort estimates, and a checklist for running it frozen while you decide.
Read moreAugust 25, 2026 · Matt Irvin
Talend MDM Server reached end of life on December 31, 2024. How to export the data models and match rules, map stewardship workflows to modern equivalents, decide whether you need an MDM platform at all, and keep golden-record logic alive in the warehouse.
Read moreAugust 25, 2026 · Matt Irvin
Field notes from large Talend migrations: how to inventory an estate, the components that don't map one-to-one, context and metadata surprises, and the parallel-run regression strategy that catches what the import wizard doesn't.
Read moreAugust 25, 2026 · Matt Irvin
Why tDBOutput row inserts into Snowflake are slow and expensive, how the Snowflake bulk components use internal stages and COPY INTO, when to push transformations down into Snowflake, and the validation queries we run after every load.
Read moreAugust 25, 2026 · Matt Irvin
An honest decision framework from a Talend shop: when warehouse pushdown and ELT beat a job engine, when row-by-row transformation and complex Java still justify Talend, and how the cost models actually compare.
Read moreAugust 25, 2026 · Matt Irvin
Log-based versus query-based change data capture, what Talend's CDC components do and don't do, when a Debezium-style pipeline is the better fit, and the warehouse merge patterns that handle late-arriving rows and soft deletes correctly.
Read moreAugust 25, 2026 · Matt Irvin
What LLMs genuinely do well for integration teams today: documenting legacy Talend jobs, accelerating migration analysis, drafting data-quality rules and schema mappings. Plus the guardrails that keep a model from silently transforming production data.
Read moreApril 2, 2014 · Matt Irvin
In this job, the tMemorizeRows component will be demonstrated so you can use it in your own applications. In this context, we will use it to check individual rows against each other, specifically to check start and end dates of items. The output will be an indicator of any information that may be in error.
Read moreMarch 20, 2014 · Matt Irvin
In some jobs, using a tMap to extract just two columns of data isn’t worth the mapping. Especially if there are operations you need to perform on the data. The tAggregateRow component can easily solve this problem, with a wide array of functions to quickly give you the information you need.
Read moreFebruary 4, 2014 · Matt Irvin
This tutorial is all about how to make sure your information is taking the proper paths in your jobs. In it, we’ll cover a bit of how tMap outputs work, as well as some of the component triggers that can happen in a job.
Read more