python - Adding a new column in Data Frame derived from other columns (Spark)

Question

Welcome To Ask or Share your Answers For Others

python - Adding a new column in Data Frame derived from other columns (Spark)

posted Oct 24, 2021 in Technique[技术] by 深蓝 (71.8m points)

python - Adding a new column in Data Frame derived from other columns (Spark)

I'm using Spark 1.3.0 and Python. I have a dataframe and I wish to add an additional column which is derived from other columns. Like this,

>>old_df.columns
[col_1, col_2, ..., col_m]

>>new_df.columns
[col_1, col_2, ..., col_m, col_n]

where

col_n = col_3 - col_4

How do I do this in PySpark?

See Question&Answers more detail:os

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

1 Reply

深蓝 · Answer 1 · 2021-10-23T19:07:31+0000

One way to achieve that is to use withColumn method:

old_df = sqlContext.createDataFrame(sc.parallelize(
    [(0, 1), (1, 3), (2, 5)]), ('col_1', 'col_2'))

new_df = old_df.withColumn('col_n', old_df.col_1 - old_df.col_2)

Alternatively you can use SQL on a registered table:

old_df.registerTempTable('old_df')
new_df = sqlContext.sql('SELECT *, col_1 - col_2 AS col_n FROM old_df')

Categories

python - Adding a new column in Data Frame derived from other columns (Spark)

python - Adding a new column in Data Frame derived from other columns (Spark)

Please log in or register to add a comment.

Please log in or register to reply this article.

1 Reply

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags