hi..I have two text files and i want to merge into one output file,also some nan values are present in this files and i want to delete from it.kindly help me.Sample files are attached here.

 採用された回答

dpb
dpb 2015 年 9 月 16 日

0 投票

Simplest here is probably just brute force...open an output file, then open each input in succession, copying lines (excepting for the first, without the header lines) from input to the output making in text substitutions need along the way.
You could, of course, read each file into memory, do the clean up there and then write a formatted file. If the input files are truly huge this may be worth doing but the other is quick 'n dirty for reasonable file sizes...
fidO=fopen('output.txt','w');
fid=fopen('input1.txt','r');
while ~feof(fid)
for i=1:2 % do header lines
l=fgets(fid);
fprintf(fidO,'%s',l);
end
l=strrep(fgets(fid),'-99.9',blanks(5)); % floating missing value
l=strrep(l,'-99',blanks(3)); % integer missing value
fprintf(fidO,'%s',l);
end
fid=fclose(fid); % done with first input
fid=fopen('input2.txt','r');
for i=1:2, fgets(fid); end % skip header lines
while ~feof(fid)
l=strrep(fgets(fid),'-99.9',blanks(5)); % floating missing value
l=strrep(l,'-99',blanks(3)); % integer missing value
fprintf(fidO,'%s',l);
end
fclose('all')
clear fid*
You'll note above for just two files I didn't even bother to create the outer loop over them; for more it would be worth doing that refinement.

5 件のコメント

skyhunt
skyhunt 2015 年 9 月 17 日
Thank you replying me, i want to remove .aws and .smd11 extensions in output file. and also some of the coloumns showed that -99.9 values again. how to correct that?
dpb
dpb 2015 年 9 月 17 日
編集済み: dpb 2015 年 9 月 17 日
Same way, make the substitution of the character string not wanted by blanks(length(unwantedstring))
For the alphabetic substitution it's more robust to use
l=strrep(lower(l),'.aws',blanks(4));
in case there is any instance where the extension is in uppercase letters.
As for the remaining numeric values, would need to see the exact code and data used that failed; only thing I can think of would be either
  1. didn't save the modified string into the variable by assigning the output from strrep, or
  2. didn't do the floating point case first or
  3. there's a case that isn't an exact match of -99.9
ADDENDUM Just as a test,
>> fid=fopen('input1.txt');
>> fgetl(fid);fgetl(fid);
>> l=strrep(fgets(fid),'-99.9',blanks(5))
l =
ZULLU.aws -04.0 -99 014 290 -99 0.1 003 3
>> l=strrep(l,'-99',blanks(3))
l =
ZULLU.aws -04.0 014 290 0.1 003 3
>> l=strrep(l,'.aws',blanks(4))
l =
ZULLU -04.0 014 290 0.1 003 3
>> fid=fclose(fid);
>>
to make sure the wrapping the fgets call inside strrep didn't somehow not do as expected.
Oh, I see one possible problem--your text file didn't have a \n character at the end of the last record; is it only the last record that is in error by chance? If that is so, solution is to either ensure the file does contain the newline after each line or, modify the while feof loop control to use
for i=1:2 % do header lines
l=fgets(fid);
fprintf(fidO,'%s',l);
end
while ischar(l)
l=strrep(fgets(fid),'-99.9',blanks(5)); % floating missing value
l=strrep(l,'-99',blanks(3)); % integer missing value
fprintf(fidO,'%s',l);
end
and similarly for second loop. NB: you'll have to return the value from fgetl there instead of throwing it away as does present code in order for it hold the proper value for the initial test the first time thru the loop.
skyhunt
skyhunt 2015 年 9 月 21 日
thanks...i have little confusions..finally -99.9 values are removed.but still .aws and .smd11 extensions appear here. please check this code.i attached script here.
dpb
dpb 2015 年 9 月 21 日
編集済み: dpb 2015 年 9 月 21 日
You didn't insert the lines to do the other substitutions into the script...I illustrated the specific line for the ".aws" extension above; your task is to incorporate that and any other patterns needed.
l=strrep(l,'.aws',blanks(4));
The general idea should be apparent by now; use as many such as above as needed to find all the patterns to be removed.
I was presuming from your example this was a pretty small number; if it's such that there are quite a large number of possible extensions so that enumerating them all is difficult or impossible, then you'll need to take the line and find the extension location within it with a search pattern and make the generic substitution or use regular expressions via regexp
skyhunt
skyhunt 2015 年 9 月 21 日
Thanks..u save me..finally i got it

サインインしてコメントする。

その他の回答 (0 件)

カテゴリ

ヘルプ センター および File Exchange で Language Support についてさらに検索

タグ

質問済み:

2015 年 9 月 16 日

コメント済み:

2015 年 9 月 21 日

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by